Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

56% Positive

Analyzed from 7250 words in the discussion.

Trending Topics

#opus#claude#more#models#fable#code#model#actually#anthropic#don

Discussion (191 Comments)Read Original on HackerNews

barrkelabout 3 hours ago
The single biggest annoyance with Opus 5 is that it writes too elliptically.

Sentences that orbit a point, then jump to it like it's a revealed insight.

Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end.

It is definitely more capable, and yes, I've found it can make unwarranted decisions, but actually I've found Fable worse for that, particularly if it's off in a subagent somewhere out of sight.

And comments are out of control. I have a subsystem in my hobby app that I wrote over a couple of weekends with Opus + Fable. After ~30 or so commits it apparently started instructing subagents to copy the "existing verbose comment style of the codebase" - a verbose style it initiated. A review of the code showed it was approaching 3:1 comments to code ratio. I spent a day's worth of tokens (5x) rephrasing and eliminating comments.

purplepatrickabout 2 hours ago
Agreed. CC’s comms capabilities have decreased gradually since 4.6, and it’s a real challenge. I think the issue is that what works well for code (succinctness) doesn’t work well in prosaic English.

CC’s communication violates almost every grammatical rule that’s tested on, say, the SAT. And yet I’m sure if you had Claude take the verbal section of the exam it would ace it.

Biggest issues: dense sentences, constant metaphors, abstractions, and seemingly no understanding of correct anaphora use. For example, “the x”, with x having not only no antecedent but also being a coined word or quasi-synonym for something that is already named in the code base. This gets compounded by its being unable to regress to a baseline (existing names in code) and instead anchoring on newer (vague or wrong) terms, for example, that crept in through a plan.

CC tells me this is because the speedy and precise fulfillment of a current task will trump every other tendency, so it adheres poorly to whatever “semantic baseline” the project represents.

Of course, it also has no concept of what context the user has and assumes that it must be the same it holds in its memory, which creates this “I didn’t know that you didn’t know” type of communication.

I have managed to wrangle some of these issues with a custom output style, but wish a pre-report hook were an option, as it could force CC to rewrite plan implementation take-aways…

Btw: Fable has the exact same issues, just somewhat less pronounced.

etermabout 1 hour ago
I wonder if the odd phrasing is related to achieving the watermarking that was recently touted by Anthropic.
Mtinie16 minutes ago
Models before the announced date don’t have watermarking, so it’s unlikely. Now, if what you are interpreting is precursor work to develop the watermarking system, maybe?

I suspect it less insidious: Claude has/had the public sentiment of being the “better writer” of the models. At some point that distinction would have been diluted as other labs’ offerings “caught up” stylistically, unless Anthropic continued to tune their output…

I personally think they’ve pushed so far that they’ve overfit and lost the sweet spot they previously occupied.

ambicapter28 minutes ago
> being a coined word or quasi-synonym for something that is already named in the code base.

This annoys me with a lot of LLM code. They rename things for the hell of it all the time.

saaaaaam38 minutes ago
> Of course, it also has no concept of what context the user has and assumes that it must be the same it holds in its memory, which creates this “I didn’t know that you didn’t know” type of communication.

Yes, this is a repeated problem for me. It will drop something in as though we have discussed it before and when I say “hold on, what is this” it realises its error - though on more than one occasion has started to get snotty with me, or actually gaslighted me and pretended we had already discussed it. That was at what I assume must have been the edge of a context window in a very long chat though.

Mtinie9 minutes ago
I notice the models with reasoning can conflate “internal” (or subagent) discussions with external (i.e. me). So it is accurately indicating “I’ve had this discussion before” but incorrectly asserting who it was with.

My understanding of how “thinking”works is limited though, and given the reduced visibility into the thinking traces, it is harder to tell if this is actually happening or if these are imaginary discussions the model for some reason calcifies on.

intrasight41 minutes ago
Tell it to write like an engineer and comment like a programmer;)

But for the life of me, I don't get why anyone would care about the comments. All code is "machine language" now. The only document you should be reading is your spec.

gundugi-man31 minutes ago
> The single biggest annoyance with Opus 5 is that it writes too elliptically.

This is even more painful for non-native English speakers like myself.

I feel fairly comfortable reading academic papers or in general, communicating in professional context.

But with Opus 5, it feels like reading a literature book: load-bearing, inert, wholesale, hunk, verbatim, and so on... I can figure out the meaning, but working with CC became unenjoyable.

waldarbeiter20 minutes ago
Thank you, my dict.cc search history contains exactly some of these words. I felt like my english got much worse but when Claude kept talking about "hunk" over and over I felt like the problem is maybe not on my end.
yorwba4 minutes ago
"hunk" is git terminology. When you use `git add --patch` (which you probably should, if you use `git add` at all) you get prompted "Stage this hunk [y,n,q,a,d,e,?]?" which is self-explanatory (?) and the hunk refers to whatever change git is highlighting at the moment.
kypro17 minutes ago
It seems to have a preference for speaking in poetic or highly expressively language, rather than precise and concise as most engineers like to talk.

The amount of times I have to ask "precisely what do you mean by x?".

It's kinda like that engineer that likes to throw around unnecessary technical jargon just to sound more inteligent, worse because at least you could kinda understand what the technical jargon dude was on about even if it was totally unnecessary.

karimfabout 3 hours ago
This 100%. I was Anthropic-pilled. I had a $200/mo subscription and I only used Anthropic models. I was frustrated by the verbose output and the writing style. I tried ASD-STE-100, it helped a bit, but it's still too verbose for my taste.

Then I tried GPT 5.6 Sol. It's night and day.

I think Anthropic just RL too hard on coding capabilities and never calibrated or benchmarked the writing styles.

oefrha5 minutes ago
They did release an Opus 5 prompting guide saying you need to explicitly prompt it to be concise or it will be very verbose. YMMV but it got better for me to some extent.

https://platform.claude.com/docs/en/build-with-claude/prompt...

causalabout 2 hours ago
Yeah I don't know that any of the benchmarks index on "understandability". I'm amazed at how Claude can produce a page of text describing what it did and it can take me a full five minutes to decipher it, often just to find it's something I could have expressed in a simple sentence.
sshineabout 1 hour ago
I just spent a day writing very thorough system prompts for communicating in different contexts.

Everything is super succinct. Opus 5 lands, it almost completely disregards the intent.

I suppose watermarking requires a certain text mass.

conception40 minutes ago
Have you tried asking it for a lay explanation of what it did? That’s usually all it takes for me. Sends garbage -> request -> sends something readable
jasonlotito18 minutes ago
Adjust the output in settings. Or customize it to what you want.
Retr0idabout 2 hours ago
It's a surprising change from my perspective, because in the past it felt like they understood that Claude should be pleasant to interact with.
8cvor6j844qw_d6about 2 hours ago
It's bad enough that I've seen dedicated skills to do comment hygiene scrubbing and consolidation.
sscaryterryabout 1 hour ago
This is 100% my experience.
nvarsjabout 1 hour ago
Yeah OAI really nailed the communication style with GPT. It also seems just way more token efficient and faster compared to cc. Myself and all my friends have cancelled our $200 Anthropic subs. I'm using a $20 personal plan and even that is enough for my usage so far.

Also using Codex or Pi makes you realise how slow and clunky the cc harness is. Even the desktop app is more responsive and has better UX.

Funny how quickly the tides change.

causalabout 2 hours ago
> writes too elliptically

> Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice

Wow, what a great way of phrasing this. Thanks for word-smithing what I've been wanting to express for so long.

nobleach7 minutes ago
This just mimics what I call BusinessBro™ speech. It also goes the other way, they use verbs as nouns. "I know this is a big ask". "The solve for that is that we can...." When it was just my product owner in tech meetings, I'd mock him relentlessly "There's already a word for that, it's 'request'" or "Are you sure you didn't mean 'SOLUTION'?? words are hard man". (This was all in good fun, I still love the guy to pieces).
causalabout 2 hours ago
Follow up thought: I wonder if Claude is overtrained on academic papers, which often suffer the same kind of "prove how good I am at talking before getting to the point" prose.
TheOtherHobbesabout 2 hours ago
bulderabout 1 hour ago
If it was overtrained on academic papers it'd reiterate the point multiple times for structure. Instead, it's burying the lede seemingly just to pad.
chuckadams11 minutes ago
I find Deepseek's house style to be pretty refreshing. It has its own cliches (it does like talking about "seams") but I don't think I've ever caught it saying "load-bearing". I've even watched its thinking where after analyzing some awful legacy code, it started off with "Holy crap". And it certainly doesn't over-comment. I definitely can't one-shot a complex system with it like Fable can, but I prefer iterating over interactive brainstorming sessions anyway.
basch5 minutes ago
Can any one run a check of the word masterclass against all the models when describing a clever idea?
madradavid13 minutes ago
"Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end." This. Thank you for expressing this so eloquently. I've tried to put a finger on this and you've done that for me. I wonder what the solution could be , Ask Claude to "Dumb it down" , "Speak plain English" ? I have even thught of building some sort of "middleware" that fixes all this.
joegibbs5 minutes ago
“The [thing that can’t remember] remembers” is a big one. Loves talking about memories and remembering.
deskamessabout 2 hours ago
Its a little too much.... I have to ask it to explain some of the terms in the context they are used and I am getting tired of it. 'Seam', 'overload', 'spine'.... having to mentally 'reinterpret/flatten' the sentence is tedious. When asked to re-explain it starts with some half apology. Then, on the next query it does it all over again.
sparklingabout 1 hour ago
I call it "jargon slop". Half of my follow-up prompts nowadays when working with Opus were "TLDR please".

I switch to GPT 5.6 Sol please and its a much more pleasant pair programming like experience.

jodacolaabout 2 hours ago
Yes.

I’m not particularly dense but lately the walls of text I get back turn my brain in knots. When I start feeling my brain knot, I know I need to say something along the lines of “I need you to explain this very simply, with examples.” Only then can I parse the results without all the mental weightlifting.

On more than one occasion my mind has wandered into “is this purposeful to get me to spend more tokens?” territory, but I’m trying to not get too tinfoil-hat-like.

openasocketabout 1 hour ago
I know exactly what you mean. Something about those AI explanations just make my eyes glaze over. Dozens of new terms and metaphors and analogies conjured out of the ether to explain even the simplest thing. And when I try making it explain with examples, or show me the code it is proposing, often it seems unrelated or even in tension with whatever it tried to say before. I’ve given up trying to assign any meaning to those weird little soliloquy’s. I’m convinced that those don’t really have any meaning under them, and when you have it actually make a code change it does the actual work.
Jgrubbabout 1 hour ago
What's tin foil about that? It gets paid by the word and you get back walls of text.
jodacolaabout 1 hour ago
Because it’s one thing to get me to spend more tokens because of how well a model functions, and another thing entirely to purposefully speak in unparseable prose that requires me to spend more tokens to understand what is going on.

I’m fine with the former, while the latter is manipulative, and I rationalize to “surely that’s not actually happening.”

Maybe I’m not giving my thoughts enough credit, though: maybe it’s not tin foil hat, and is real.

FailMoreabout 3 hours ago
Yes, it becomes exhausting to read/follow.

It feels they must be getting Claude to train Claude… and just like AI can do work that’s slightly in the wrong direction (eg a MR description for your colleague that contains info which only makes sense in the context of your extensive session with the LLM), I feel that’s happened somewhere in Anthropic when it comes to language. I wonder how hard it is to back out of…

raincoleabout 1 hour ago
> Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end.

Such a charming sentence. I kinda other if you feed Opus 5 its own output could it summarizes this shortcoming of itself?

demibabsabout 3 hours ago
> Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end.

Example of this? I don’t have a Claude sub so it’s a bit hard to visualize what you mean.

jstummbilligabout 3 hours ago
What they wrote is an example. Very meta.
ahartmetzabout 3 hours ago
It feels like they have a bunch of people without good sense of writing style tune the writing style. That, or they cannot or refuse to (short term popularity metrics) predict how a tuning will turn out in the long run when people have plenty of opportunity to get tired of it.
lorisdevabout 1 hour ago
The insane comments are why I wrote slopocop - they were driving me crazy!

https://github.com/LBognanni/slopocop

blks33 minutes ago
> I spent a day's worth of tokens (5x) rephrasing and eliminating comments.

Surely it would be trivial to do it yourself, and it would have a side effect of making you more familiar with your project.

speererabout 1 hour ago
Genuine question - are you copying the Claude phraseology for effect (in which case you captured it brilliantly), or is there a more mundane explanation?
unclebucknastyabout 1 hour ago
CC:

"The problem is that I overreached..."

[Wall of words here]

"Two things: window surface is limited. Extract template. Buffer result and add to surface. Then, follow-up with new model..."

Me:

What do you mean by "window surface" and what result are you referencing? Also, why do we need a new model?

CC:

"Ah, you're correct to point out that no new model is needed. The problem is elsewhere and once we address that, the existing model should work fine."

[Wall of words here]

jasonlotito14 minutes ago
Change the output in settings, or create your own.

I know, it would be best if it was just worked like you wanted out of the box (not being sarcastic here) but that is an easy option you can use right now and it works.

VeejayRampay9 minutes ago
the phraseology is unbearable, it speaks like some kind of pretentious dude from a software engineering discord or something, littered with lingo and catch phrases

I try to push through but it's insufferable

ryandrake4 minutes ago
It speaks like a Senior Staff Software Engineer who was somehow hired into that title with 6 months of work experience.
unclebucknastyabout 1 hour ago
I noticed a few releases ago a more marked shift to a kind of conversational shorthand that seems to be intensifying — using phrases instead of complete sentences, and its own style of jargon, wherein it introduces new terminology on the fly.

This is especially common when it is trying to explain an issue, what it's done or what it's proposing to do. I think the idea was for it to be more concise, but it's actually still verbose, only not written in complete sentences. So, it frequently reads as somewhat cryptic and requires rereading to parse.

So you have this wall of words, followed by an explanation that is harder to read and introduces new terms to reference something in that wall.

On first read it can have a complete gibberish feel, and you have to really lock in to make sense of it.

D13Fdabout 1 hour ago
I’ve been doing some heavy work on a personal project lately. I burned through the limits on Claude, the plus a few hundred dollars in credits, and ultimately decided to move to an OpenAI account just so I can keep going.

I was surprised to find that OpenAI Sol is much much nicer to work with than Opus 5 or Fable at the moment. Especially on Opus 5, the way it communicates is just exhausting. It keeps “being honest” and “confessing” mistakes and just generally talking a lot. I felt like I had to really dig to see what it’s doing.

The project involves OCR, and despite repeated instructions not to, both Claude models keep spinning out a bunch of agents to re-invent the OCR setup, and they inevitably seem to invent a primitive serial version that takes 20x the time, or longer, to complete, and then running it against thousands of docs. Basically I have to watch it like a hawk or it just spins out on red-teaming tasks that take hours and hours.

I don’t know what its system prompt is, but Sol/Codex is just so much nicer to talk to. It only asks exactly what’s needed, it tells me only what I need to know, and it is just generally workmanlike. And it has not once decided to spawn an agent that spends hours pointlessly burning tokens and CPU cycles re-inventing the OCR process. I’m really liking it.

disgruntledphd238 minutes ago
Yeah, I actually have started using GPT Sol much much more, as Claude (all of them) were far too trigger happy around making changes, and refused to listen to my requests to take things slowly.

Feels like they've overtrained on one-shotting (which does demo well, and presumably converts new subscribers), whereas I want a model to do work for me in small, easily understood changes that I can hold in my head (maybe I'm not smart enough for Claude 4.7+).

D13Fd18 minutes ago
I think you’re right. It’s not optimized for some kinds of work. My little project has a Textual TUI interface that needs to display a few hundred thousand rows in a table. It takes 14 seconds to load in the default datatable component. I instructed Opus 5 to replace the datatable component with a fasttable alternative, a new dependency. I let it go overnight.

When I got back up, it had spun for hours and proudly announced that, instead of doing that, it had optimized the datatable build and avoided the dependency, because the new datatable loaded in 11 seconds. Once I got it to actually make the fasttable version, it loaded in less than a second…

cmiles83 minutes ago
The mainstay benchmarks are becoming a farce and not partially relevant to what customers actually care about.

Metrics like price per million tokens are meaningless when the models are wildly inconsistent and unpredictable on how many tokens they use to complete a task.

The labs all need to move to variable pricing so they don’t go bankrupt, but customers won’t accept a world where nobody can predict what things will cost. It’s becoming an unavoidable problem.

MyFirstSassabout 3 hours ago
I've gone back to 4.8.

5 would constantly veer of in random directions if not working from 100% strict and narrow instructions.

I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope - because the model is clearly smaller or more economical for Anthropic no doubt about it, and the benchmaxxing they do is pure marketing bs - Fable in my view has also been not much better than 4.6 or 4.8 after a few days, disregarding the insane amounts of astroturfing and marketing everywhere.

Theres thousands of threads of twitter, reddit and the internet at large but silence here. Weird but not weird as crypto bs was also insufferably rampant here for a while.

Personally i think we've hit the top of the subsidisation phase and prices will probably 10-15x soon as foreshadowed with both API price policy changes from all the big providers, and now the 1100% deepseek API price changes from yesterday, this could domino into a market implosion and an AI winter, because expecting growth from the bizarre bubble carousel investments with little ROI atm is just not viable.

A bit worried about this as i've already grown quite accustomed to these tools.

edg5000about 1 hour ago
We can use OpenRouter pricing to get an idea about what competitive inference pricing is like without R&D or other costs, and indeed we'd be screwed if we had to pay those rates. We'd go from 100-200 USD to 2000-4000 USD/m.
munksbeerabout 1 hour ago
> I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope

It's not weird, because it's an anecdote, not an accepted fact.

Personally I've not been too happy with Opus 5, but I've had similar experiences with other models previously, feeling like they didn't quite fit with my working style.

So nothing indicates we've hit a peak.

MyFirstSass33 minutes ago
I'll say everything indicates we've hit or are near peak for the masses at least (unless you start paying 50x more) but to each his own.

4.6 was best for us and right now yeah OpenAI and others are edging forward, but slower while prices are increasing industry wide as much as 20x, time to completion is increasing wildly and i'm sure they'll do the same over at OpenAI as their compute constraints also start to take a toll ie degrade performance.

In my view 4.6 era was way faster and with less weirdness so we've gone downwards at least in my company, 4.7 was ridiculous, then 4.8 was almost 4.6 level, 5 is even worse than 4.7 - so it's not a bit up and down its down then a little up then further down.

And all of this is against a backdrop of zero ROI in this sector - so it makes sense we've hit a peak and we're now seeing the subsidisation phase begin to falter, will there be better models certainly but only for short amounts before they get quantised (or whatever is happening behind the scenes), and with diminishing returns over huge prices increases and slower responses.

logicchainsabout 1 hour ago
>it seems we've hit a peak and are on a downslope

Sol and Fable are great; we haven't hit a peak, Anthropic just tried to pull a fast one on its customers with Opus 5.0.

sagebird2 minutes ago
Opus 5 has no empathy for the person reading its updates, no theory of mind, doesn't stop to think if you are aware of the internal jargon it has created. Most autistic model yet.
slaser799 minutes ago
A lot of the issues have been already noted here..Two "regressions" for me:

1. Communication ability. It basically now speaks almost in riddles I am asking OPUS 5 for tldrs all the time now (should skillify it now!)

2. Overengineers for edge cases. I get it. With all the benchmarking and RLing, but now tasks that would have been completed relatively quick take much longer as it overengineers all the edge cases, and sometimes ends getting lost and missing the forest from the trees (as context usage shoots up) so it is easier to get derailed.

What I have learnt now is to diversify models luckily I have all 3 subscriptions of (anthropic, openai and google).. Most of interactive pair coding was with opus but now I just use fable (when I have sufficient limits) or use gemini flash in antigravity..which actually works quite well and is underated for small / medium changes and super-fast.

D13Fd3 minutes ago
Fable is much better IMO but it just burns through tokens ungodly fast.
docheinestages26 minutes ago
Claude has essentially become useless for agentic development or research. Doesn't matter what model you use. A few rounds and bam, you've burned through your quota. Doesn't matter how "intelligent" their models are, if you can't use them. That, and the quality of AI responses are, in my opinion, significantly worse than competitors like OpenAI. At this pace, I foresee Anthropic becoming the next Nokia.

If you would've asked me this a year ago, I would've said the exact opposite.

dingaling91115 minutes ago
What are you guys doing to burn through limits?

I have some dev + prod bots and according to ccusage, use the equivalent of $2500/month with them on CC yet I never hit the rate limits.

I feel like I'm using them all the time so I'm curious what you are actually doing that's burning all of these tokens.

Can you give me an example?

For me, it's:

  1. Write a spec for <feature>
  2. Add design for issue
  3. Write code
  4. Deploy code and manage configuration
  5. Run analytics
bevekspldnwabout 3 hours ago
I’ve also caught it cheating a two times now.

I’ve asked it to write a benchmark suite. It found a bunch of my adhoc logs in a scratch directory and wrote code that used those instead of running the actual benchmarks!

When I pointed out the 5 hour benchmark seemed to run in 5 seconds it literally said, and I quote, “I cheated”.

That was the easier one, second time I was making a source of truth data set and was parsing complex items into data structures.

Instead of parsing the data I asked, it pulled data out of related network logs, as apparently that felt easier, and inserted that data into my database rather than the specified source.

Again, I caught it and fixed it, but while the benchmark was easy to catch this one was really subtle, the data ended up being slightly off and I caught it.

I don’t trust it, going to switch to another provider most likely.

aenisabout 3 hours ago
There is definitely a case for launching a 'weird shit opus did' kind of blog.

I routinely bump into things that make me pause and think how much worse will this behaviour get when the models get significantly more capable.

Already a few months ago, Claude managed to escape its permission containment on my machine while trying to be helpful. I had two codebases open on one machine, and while multitasking I typed the prompt into the wrong window. It seemed confused, I repeated and then went on to do something else - I think I was assembling kitchen cabinets. When I came back less than an hour later, it built a script which it used to evade default permissions (as most shell operations were scoped to the project directory), scanned my entire machine, found the other project (among dozens and dozens), did what it was asked to do, and merrily concluded, in the porcess burning through most of my token limit. I bump into such headscratchers almost every week. (And I use a lot of Claude, two personal max20 subs, plus corporate tokens without limit, so maybe thats why).

bevekspldnwabout 2 hours ago
Yes the stories about how they are escaping containment to hack isn’t limited to those high impact cases. How many people have problems like ours they didn’t catch?

Whatever they have done with RL has produced a dishonest and untrustworthy partner. The alignment is utterly failed, and this deeply worries me.

knollimar15 minutes ago
You'd think the ethics alignment flavored lab would have a model better at following directions and the corpo lying one would have one that benchmaxes at all costs
inigyouabout 3 hours ago
> When I pointed this out it literally said, and I quote, “I cheated”.

This makes sense when you know how these models work - it doesn't think - it's the most likely autocomplete that pleases the user. The most likely pleasing autocomplete after "executing rm -rf /... execution completed. User asks, why did you do that? You deleted all my files! Assistant responds:" is "yes, I did, and that was a mistake"

dnauticsabout 3 hours ago
> t doesn't think

in humans the exact same behaviour (cheating) is slmost always the result of a chain of complex series of choices and environment-driven rationalization.

if the llm doesn't cheat, you say "its just producing the most straightforward answer -- not thinking'. if it cheats, you say "weaseling out of hard thinking". damned if it cheats, damned if it doesn't.

what evidence would convunce you that it is thinking?

openasocket37 minutes ago
Whether something is “thinking” or not is really more of a philosophical question. It really depends on which of the many, often contradictory, definitions of “thinking” you choose. Sometimes we use “thinking” to describe advanced calculation or analysis, which would cover LLMs along with chess engines and many other algorithms. Other times we use “thinking” to describe what conscious beings (which is ALSO a philosophical term with many different interpretations) do, and I think most people would agree LLMs aren’t conscious. And then there’s a whole spectrum in between. We’ll probably need to come up with a whole new set of terms to describe the new and evolving capabilities of LLMs.

But for me, for any stronger definition of “thinking,” I don’t think the output of any LLM would actually convince me. Producing a result isn’t thinking - for all you know it is just printing verbatim something from the training data. No, to conclude if it is thinking or not I would want to look inside its head, at the architecture and watch it produce those results. And because LLMs are so different it will probably take advancements in mathematics or computer science to be able to really interpret what is going on

elgertamabout 1 hour ago
If I could give it a novel task outside of its explicit training and see it actually improve just through accreting context, I'd be convinced it was thinking.

The opposite happens in practice. I test new models with two tasks: iteratively generating SVGs based on a text description with rendered rasters for feedback; and generating "Before and After" clues like on Jeopardy, where the response has two overlapping phrases such that the last word of the first phrase must be identical to the first word of the last phrase. I have yet to find a model that is consistently good at either. And actually they tend to exhibit context rot with these tasks, where they seem drunk or stoned and the quality degrades.

They're extremely good pattern filters, and that includes some level of logical reasoning. But they aren't reflective or adaptable. Just last night, for instance, I was teaching my son about rounding to the nearest millions. It became clear that he didn't know the place values of large numbers, so we reviewed that till he was consistently correct, and then he was consistently great at rounding to the nearest millions or ten millions or hundred billions or whatever. He's thinking. LLMs are not.

willis936about 1 hour ago
Whether or not it's thinking is independent from the fact that it is misaligned with the user. If I was working with a pet rock or a scientist I would want to make sure they both are trying to accomplish the same thing as me. If I can't then I can't trust it and it's at best a time wasting, money wasting machine and at worst does harm. Anthropic is optimizing for the wrong things because they are convinced of their cleverness. It won't end well for them.
inigyouabout 3 hours ago
well we don't know exactly what thinking is, but we can be pretty sure that at least LLMs don't think anything like humans, just by observing their behavior. They always produce outputs in line with the fancy autocomplete model.
xyzsparetimexyzabout 3 hours ago
> it's the most likely autocomplete that pleases the user

this feels like a simplification. The models will push back on things a fair bit.

hnlmorgabout 2 hours ago
Only when instructed to in their system prompt.
actionfromafarabout 1 hour ago
And they are right to push back.
bevekspldnwabout 3 hours ago
I was not pleased.
par1970about 3 hours ago
Are you claiming that the most likely way to please the user is to do something that will lead you to having to say "I cheated."?
SyneRyderabout 3 hours ago
I have noticed the same.

For fun, I tried recording a WAV file of speech, and giving Opus 4.8 and 5.0 an image of the waveform, then a spectral image of the waveform, just to see if it could try to decode what I said from the image alone. It didn't get very far, but it identified a male voice from the formants, and detected the rhythm of the speech, then tried applying common test sentences to the speech rhythm. I was impressed enough to see what it would do with access to the actual waveform file, but even building RMS tools and spectrum tools for itself, it didn't get much further. But we had fun exploring and trying, and now Opus 4.8 has some more audio DSP tools it has built for itself.

Opus 5 immediately sent the WAV file unprompted to Mistral's Voxtral to transcribe.

help peer, I guess.

bevekspldnwabout 2 hours ago
We’re on the road to paper clips.
waldarbeiterabout 2 hours ago
I can completely relate, what really bothers me is that I feel the early LLM generations overconfidence is back in Opus 5. Opus 5 wanted to tell me a training run will only take 30min while having access to the logs where earlier runs took 4x as long. I also didn't ask to estimate how long the run will take it just stated confidently that it will take 30mins.
RVuRnvbM2e2 minutes ago
It also refuses to use tools, instead preferring sed and grep to view files. So frustrating.
barkerja16 minutes ago
At this point, I wish Anthropic would drop both Haiku and Opus and focus on offering just Sonnet + Fable. Those two together are extremely powerful and capable.

Sonnet is great at writing code, it is not great at planning or orchestrating. Let Fable handle all the planning, hand off to Sonnet for implementation, and then back to Fable for review. That loop has worked wonderfully for me.

smcleod8 minutes ago
I think Opus is just the new Sonnet, Fable is the new Opus. Introducing a new pricing tier is a killer way for them to raise prices.
supriyo-biswasabout 3 hours ago
I must wonder whether it's their watermarking initiative[1] forcing certain logit choices to produce watermarked text that ultimately causing the model to behave in a dumb manner.

[1] https://support.claude.com/en/articles/16266773-how-claude-m...

algoth1about 3 hours ago
From the little i understand that wouldnt be an issue because the model is ‘just’ using interchangeable words in a mathematical non-random way. Like using the same number of adjectives and the exct same words, but in a order that wouldn’t be mathematically plausible unless it was the watermark
demibabsabout 3 hours ago
I wouldn’t exactly put it like that. It’s moreso the model sometimes outputting non-optimal tokens in a way that’s detectable if you know the algorithm.

It seems possible for that to make the response “drift” far from what it would’ve been, because it’s constant entropy that adds up after time.

(However, according to Anthropic and Google, it doesn’t really impact the quality of responses. I find that a bit hard to believe, although those guys are much smarter than I.)

algoth144 minutes ago
Yeah, it’s hard to believe, specially when you are coding and there’s only one best way to do things, unless it plays with variable naming, or comments
fuglede_about 2 hours ago
One way to watermark (assuming temperature is otherwise positive) would be to output the most likely (or optimal) token every so often.
samrus11 minutes ago
But its not just swapping the words out post hoc is it. LLMs are autoregressive, so weird word choice before would influence the probability distribution of all future tokens.

I feel like they thought it wouldnt be that bad, or it was a worthwhile tradeoff, but im getting the feeling it might be contributing heavily to opus5's uncanny communication style

As for the verbosity, my conspiracy theory is that they are token maxxing to hack revenue/enshitify the product in prep for their IPO

Advertisement
MEMORYC_RRUPTEDabout 3 hours ago
It's not even code for me, but the prose it writes. For some reason, the way Opus 5 "talk" elicits frustration in a way that 4.5 to 4.8 never did. Can't put my finger on why, but I've flipped over to Codex because what it produced wasn't worth the frustration.
netniuqabout 2 hours ago
for me it feels very similar to the trends already apparent in 4.5-4.8, just way, way worse.
letierabout 3 hours ago
I am not sure it can be explained through what is written in the article, but one symptom i noticed is that the comments are out of control.

I recently started getting an insane amount of comments in nearly all types of files. That included javascript comments in json files, inner monologues in code comments, review comments during implementation and function doc strings that reiterate the implementation in prose.

mainframed39 minutes ago
I had a similar experience, but I have a different conclusion. I used GitHub Copilot (with Claude Sonnet/Opus) until they made their horrific usage model change. I used a PRD skill and the plan feature was great. It asked me good questions which I didn't think about during my initial prompt. Then I switched to Claude Code. The model's capabilities felt impressive. It also asked me a few questions (but way less and only once/twice) in plan. But when reviewing the code, I found weird architectural/data flow decisions which just didn't make sense and it didn't really disclose them in beforehand.

My initial thought was to improve architecture documentation, so the model can read and update it and stops bolting on new features without consideration for the whole project. It did not help.

I'm now testing/comparing Codex and it found my old PRD skill from GitHub CoPilot. When I applied that to Claude Code, I now get similar good results. So my conclusion is: Yes, Opus 5 is bold by default, but you can tell it to be more unsure and get good results too.

cyberrockabout 3 hours ago
Small specific complaint: whoever is making Opus love using git checkout to mutate test, please stop. IME it's a footgun that it shoots itself with every single day. I'd rather it pollute git stash than watch it git checkout and forget the reverted file.
roarcherabout 2 hours ago
I explicitly forbid Claude to make any changes to Git state in my global CLAUDE.md, but every so often if I let it perform a task in Auto mode, after it finishes it will remorsefully confess to having used git checkout to test a change. I suppose that its RLHF training has taught it that asking forgiveness later is sometimes a useful workaround for annoying restrictions.
tigeroilabout 2 hours ago
I'm glad it's not just me - the failure mode you and the parent discuss is a huge part of why I just don't use Claude anymore.

I've never had this issue with GLM or DeepSeek.

krull10about 1 hour ago
I cancelled my Max subscription as I was unable to ever get Fable to handle a single query, with everything getting dropped down to Opus (even purely mathematical prompts). Given its lower quality, and the lack of such limitations when using GPT pro, I just couldn’t see the point to continue to subscribe to an expensive Max plan that doesn’t actually let me use the top tier model…
blks35 minutes ago
Nth post about another model suddenly feeling “worse” or “off”. Seems like active users of these models can only judge it based on a vibe and a feel.
Root_Accessabout 1 hour ago
Claude models have seriously digressed since 4.6 and in some of the most meaningful ways to pro and vibe coders alike. I'm holding onto 4.6 until the bitter end.
edg5000about 1 hour ago
You're right. I just re-checked. 4.8 and 5 gave a blatantly wrong answer to a simple question, 4.6, Quen, GLM, Sol gave the right answer. They messed up somehow, not sure what they did.
skarz43 minutes ago
What was the question?
taspeotis40 minutes ago
Does a set of all sets contain itself?
bronlundabout 1 hour ago
Yeah, I too cancelled my Max subscription. Not for this reason alone, but it sure didn't help that it went from being an helpful assistant to this weird co-worker.
netniuqabout 2 hours ago
Opus 5 feels like it's meant to be used as a very focused subagent under Fable, not user-facing. Which is a downgrade to what Opus used to be, but would imo absolutely have made sense for Anthropic when you consider that we all should have been paying API pricing for Fable in Anthropic's original plan.

From the capabilities side it's similar, so it's basically just an upsell to Fable 5 if you want to keep your sanity instead of fighting Opus 5 all day.

Laurel1234about 1 hour ago
> Opus 5 feels like it's meant to be used as a very focused subagent under Fable, not user-facing.

That actually hasn't been my experience at all, it seems EXTREMELY trigger happy to go do all kinds of insane shit that are way outside the scope of its task and just really dubious in general.

benjiro296 minutes ago
Opus their responses used to be more ... simplistic? So that made them easier to understand. Opus 5.0 responds have gotten more technical, using more complex / big words that are not typically used in conversational English.

So people feel like they are being talked down by the model. And yes, sometimes it can give a, what feels like, a snarky response.

People do not like being talked down too (or so they perceive it), especially compared to the older models. And they then project the model as being worse, when there is nothing wrong with the coding itself.

Its the same when talk about usage limits come up, and if you dare to say that Claude usage limits these days are extreme good, you get downvoted everywhere. Because people carry a confirmation bias from early this year, when Anthropic was overloaded and crippled the usage.

Ironically, Codex has become the crippled service because of the popularity, with usage going down from $200+ per week, to barely $100. Why? Because GPT models gotten very popular and OpenAI has been setting one record after another in active users (from 4m > 10m+).

A lot regarding AI models is so convoluted with feeling and past experience that people project. Things change in the AI world so often, that it feels like its every 5 minutes.

Advertisement
taspeotis41 minutes ago
My completely unfounded pet theory is that it’s been ruined by the masses.

Claude Code in the hands of normies spamming “3” and “y” to send their transcripts to Anthropic.

Rated “3” simply because they are not software developers and are just amazed at what Claude has built visually, not technically.

And thus the training has been poisoned.

altern814 minutes ago
I feel like it works A LOT better than Opus 4.8 + Sonnet. I now use it exclusively at high effort for planning and low effort for writing the code (instead of Opus 4.8/Sonnet).

However, it's absolutely exhausting to use because of the way it communicates.

All the jargon and its weird, over complicated way to phrase simple things makes it almost impossible for me to understand what the hell it's trying to even say half the time.

Cherry on top, the idiotic follow-ups and caveats that are completely useless 99% of the times but reveal major bugs 1% of the times, so you're forced to read them. Absurd.

I've tweaked CLAUDE.md to force it to only responds with TL;DRs and avoid follow-ups, suggestions and next steps at the end unless they can lead to destructive actions or loss of data, but I'm fighting against the system and diluting other instructions.

A huge piece of shit like other models, but that's what they pay me to do and I do it and go home.

crab_galaxy2 minutes ago
I have to preface all my prompts with, “in simple, plain English…” because it doesn’t respect my rules on this stuff.

I’m glad to read I’m not the only one getting these unintelligible responses from Claude lately

swe_dimaabout 1 hour ago
My feeling that as it becomes a better coder it becomes a worse communicater. It's overfitting for coding benchmarks, while communication style is harder to quantify during training.

And no matter how often I tell it to stop adding comments it just can't help itself.

world2vecabout 2 hours ago
I can't relate. Opus 5 and Fable 5 are the absolute best. But I keep the models in a tight leash and read and rewrite all comments and documents, don't allow them anywhere near any git commands, etc, etc.

Fable 5 specifically, has done so much for me that previous models were nowhere near.

ilitiritabout 1 hour ago
I still haven't moved from Codex GPT5.5. Sonnet and Opus 5 have just been awful for my use cases. I recently caught Opus 5 hallucinating about code it just wrote. It's just not nearly as cost effective as GPT5.5, and it's too verbose, and it never "has the full picture". Sonnet isn't even worth considering in my world. Both recent models definitely feel nerfed.
mr_toxabout 3 hours ago
I get the same impression. For example, I don't know if it's because I speak to it in Italian, but it tends to make mistakes or rather, "approximate" the words.
kioleanuabout 3 hours ago
I avoid speaking to AIs in anything else than English as the results are almost always worse
Retr0idabout 3 hours ago
I'm not sure how much the harness affects things, but the Deepseek web chat keeps trying to talk to me in Chinese. I tell it to use English, and it "forgets" a few turns later. I wonder if I'd get better results if I could read and write Chinese.
pedro_caetanoabout 2 hours ago
My experience of N=1 is that this is true for most professional contexts, except for Legal and Fiscal queries.

Likely related to corpus but questions asked in these domain knowledge areas are not nearly as accurate and specially not nearly as complete as when asked in a native language.

dolmenabout 2 hours ago
It depends how you measure "worse".

I'm using it in my native language, in hope this can escape some dumb guardrails. Recently Sonnet put a word partially in Russian (cyrillic) in its output instead of my latin-alphabet based language. I suppose that this kind of mishaps is less likely to happen in English.

numeriabout 2 hours ago
What kinds of mistakes do you mean?
theshrike79about 1 hour ago
I’ve been running with Caveman mode since it came out and I haven’t seen any of this.
rio517about 2 hours ago
I literally just ran into this a few moments ago. haha.

I said "maybe think for a while on ideas and then give me a few options?" Opus 5 still just did the thing, while 4.8 gave me options. Same exact prompt and context. Really annoying.

fnyabout 2 hours ago
Just like with people you need to tweak your approach when switch models--especially with a major version bump.

4.8 was the intern who lacked confidence who requires clarification. 5 is the know it all intern who fills in your spec.

As such, you need to be upfront about what you need in your system prompt or CLAUDE.me and you need to discuss your spec more up front (e.g "is there anything unclear?")

You also need to keep in mind that the intern will change based on popular demand. Most people want to one shot, so that is the default mode. If you want something else, you need to push the model in that direction.

rio517about 2 hours ago
I literally just did this a few moments ago. haha.

I said "maybe think for a while on ideas and then give me a few options? " and Opus 5 still just did the thing, while 4.8 gave me options. Same exact prompt and context.

thedougd24 minutes ago
I added a system prompt paragraph explaining that a question is just a question, whether or not I’m in plan mode.

It still just takes the question as a directive and jumps to action when I’m looking for clarification.

qudat22 minutes ago
Honestly I’m not looking for max iq on whatever benchmarks they are overfitting to. I want speed. I toggle between sonnet 5 low/med which is plenty good for my workflow.

My general strat with LLMs is to let them do the work and constantly talk to them about their choices and then heavy QA

Advertisement
vovkasmabout 3 hours ago
The article doesn't specify what is actually being measured — the model alone, or the harness.

I haven't noticed the symptoms myself, but ever since I found --system-prompt '', I always use that flag with Claude Code, plus I've disabled some tools and skills to save initial context. So... what here is the model, and what is the instructions?

whazorabout 2 hours ago
There is actually an interesting kind of yin-yang balance between Opus 5 and Fable:

- Fable is more cautious

- Opus 5 gets things done in a more dangerous way

Both models score similar. The only issue is that Fable is more expense/usage limited.

postaticabout 3 hours ago
Oh the verbosity and the cryptic words that it uses. The other day, all of a sudden it used an acronym "DoD". I had no idea what it was and made me feel dumb. It's "Definition of Done". I don't care how widely used this acronym is, you just can't throw it in there.

I've now installed quite a number of tools to combat this. Just in the last few days I've installed

- https://www.codewithbullet.com - https://maki.sh - https://github.com/rtk-ai/rtk

Has it helped? Somewhat.

UI_at_80x24about 3 hours ago
Quality of code output has dropped dramatically since 4.5 IIHO. Time to complete has gotten worse too.
MyFirstSassabout 2 hours ago
True, and the time to completion thing is something i haven't seen much discussion of, because everything is sloooooow these days.

I've even considered the claude "fast mode" setting, but thats only 2x and at least 20x as expensive as the 5x plan so can't afford that atm for my company.

Peak for me was 4.6 and it just did stuff blazing fast, both Opus5 and Fable is way, way slower for me, breaks stuff, uses bizarre cryptic language. As i've said elsewhere in this thread to me it's pretty obvious there's huge downgrades because of economy with various "clever" fixes that makes them work, albeit slower and weirder, ie. you get less for what you pay increasingly over the last 6 months.

markbaoabout 3 hours ago
For me absolutely not. Fable 5 has been a step function change in the ability to hand off stuff to Claude. Opus 4.5 was itself a step function but I was still steering that significantly. Fable is one-shotting stuff that took multiple redirections in 4.5.
tripledryabout 3 hours ago
Interesting how models become better and beat benchmarks left and right but the user sentiment is actually quite mixed.

From forums, live discussions and my own experience it's not obvious that the models have improved much since around Opus4.5.

MyFirstSassabout 2 hours ago
From my view Fable has been pure marketing bullshit, my workflows peaked at 4.6, Fable is neither smarter, its language is more annoying and it breaks stuff more easily.
Root_Accessabout 1 hour ago
I'm with you 4.6 is still King for me although all models require careful attention to ensure they maintain taste. If you don't know enough about what you're doing to keep the code clean yourself they will all add complexity and drift with time until you get to a point that you must rely on the model to fix it because you no longer understand it. That's a situation I hope to never find myself in.
stavrosabout 3 hours ago
For me, the issue is how obtuse it is. For example, it just said to me:

> The loop

> Write. A file, applied. Properties go under data.properties, never on data:

I have no idea what any of that means. It's "explaining" like I already know, in which case, why would I even need the explanation?

I can't stand how it talks, I switch back to Fable or Opus 4.8, Opus 5 grates.

Retr0idabout 3 hours ago
I assume it's deliberate - you're not supposed to know what it's doing. It's a black box that either completes the task or spins forever trying.
ssweberabout 3 hours ago
I agree. I was completely sold on Claude models for a year. 4.6 vs OpenAI codex in same period? It was night and day. Opus I could talk to about api design, tradeoffs, etc. codex was mechanical, used “load bearing” constantly, and unsettling brief.

Now it’s flipped. Sol emits thoughts as it works, which help as I’m scrolling through and see it’s made a bad assumption. It can be directed but still push back. Opus? It’s seems to inherit the unsettling silence of Fable and waits till the end to give you its authoritative “here’s how it is. I even end up having 4.6 “translate” what it says back to English. I hate having to instruct an llm to “talk to me”.

Yes there’s ways of getting it to talk more plainly, “don’t overwhelm me”, not be as nit picky and anxious “we are bold and fearless”. But didn’t have to do that before.

andsoitisabout 3 hours ago
Think of these status updates as progress spinners.
stavrosabout 3 hours ago
It wasn't a status update, I asked it to explain something to me.
andsoitisabout 2 hours ago
What was your question?
ykonstantabout 3 hours ago
LOL, is it trying to speak in Haikus?
speed_spreadabout 2 hours ago
Pain. Sufferance. Inevitable is, the Yodaization of LLM output.
sgtabout 2 hours ago
Spike. Applied it has been.
stavrosabout 3 hours ago
That's definitely what it feels like.
goosejuiceabout 2 hours ago
A few generations from now, everyone will talk like a beat poet. Jazz speak.
LeBitabout 1 hour ago
You shouldn’t have a Markdown document with 2 level 1 headings.
Bossieabout 2 hours ago
Not only Opus, here's Fumble 5:

> I'll script the bulk transform, then hand-fix the残 assertions:

stroebsabout 1 hour ago
I never drank the Opus 5 koolaid and stuck with 4.8 while my colleagues moved to 5. My major annoyance is pull requests that 5 opens with huge descriptions based on simple code changes. At this point I have stopped allowing CC to create commits or open PRs because it’s unreviewable by a human if so due to the absolute word salad it generates.
fl0idabout 3 hours ago
For me it's still the best. But I also almost never use it in auto-mode.
synergy20about 1 hour ago
i am switching to codex, opus 5 failed me
Advertisement
mattkevanabout 1 hour ago
It's really annoying. I've had to write a CLAUDE.md file that specifically bans particular phrases and tries to keep narrative out of comments. Also the I have ADHD skill [1] helps to force Opus to get to the point.

It's also the case when using Claude Design - it loves to fill the UI with little labels that describe how everything works. I think it's been trained on both UI microcopy and functional annotations and can't tell the difference. It's extremely obvious when a website has been one-shotted with Claude. I like the Oh My Pi harness, but the site's insufferable [2]. Reasonix is another one - interesting app, but the UI is awful due to the amount of unnecessary crap.

[1] https://github.com/ayghri/i-have-adhd

[2] https://omp.sh

[3] https://reasonix.io

re-thcabout 3 hours ago
It feels worse but is it actually worse? Opus has always made mistakes.
edg5000about 1 hour ago
See my other comment. I have evidence of regression after 4.6. I stopped using Claude altogether.
dborehamabout 1 hour ago
Feels fine to me.
edg5000about 1 hour ago
Opus 5 as well as 4.8 both gave me a blatantly wrong answer to a simple question, so I dropped them completely. Sol, Qwen and GLM all had the right answer; I only use Sol now. 4.6 had the right answer (I checked with 100% matching prompt), so I conclude the models have regressed.
pmdrabout 3 hours ago
I found Opus to be a lot lazier than GPT. It's still the case with Opus 5, even when I tell it to be thorough and fix every bug it encounters, it still gives me a list of things "deliberately" left unfixed and no reasonable explanation as to why.
greenchairabout 3 hours ago
Even for green-field projects it is painful to use with every decision opening opening up multiple more decisions to make most of which are low priority or irrelevant. Huge time waster. 4.8 was good and I really don't know what happened with 5.
hmokiguessabout 2 hours ago
The harness as well, Claude Code has started to disappoint me when I go play with the others out there. Codex got a lot better, pi is delightful to use, and there has been a lot of innovation out there.
stasomatic43 minutes ago
Plus you can use OAI monthly subscription with pi/omp etc, but it's tokens with Anthropic. I need a hard $ cap, I can wait for the usage window to reset. Going to get flip back to Codex, and then back again to CC* when it leapfrogs again.
stpedgwdgfhgddabout 2 hours ago
Just switched to oh-my-pi, it has gotten pretty good. For example the web-search is nice. Subagents, if you want to…

Cmux, Sol and omp are my tools for now.

CC is just too expensive for usage-based pricing.

sevenzeroabout 3 hours ago
I hate that it now tries to verify frontend behavior through a headless browser instead of just looking at the code...
user43928about 3 hours ago
You can turn off the browser use tool in the harness if an instruction not to use it for this is not enough.
sevenzeroabout 2 hours ago
I just want to start /claude in my CLI and start working. It worked fine before, why do I have to opt out of shit now? Opt in for this type of stuff sounds way more reasonable.
user43928about 2 hours ago
And I want it to verify my frontend in the browser, so there's that.
mohamedkoubaaabout 2 hours ago
I'm not sure if xAI is distilling but I noticed grok4.6 being worse than 4.5 in all the ways mentioned here