DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
67% Positive
Analyzed from 9673 words in the discussion.
Trending Topics
#opus#fable#claude#models#model#more#better#anthropic#don#astra

Discussion (568 Comments)Read Original on HackerNews
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.
Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.
It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3
sir, this is a hacker news thread
They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.
Do you have a better proposed phrase?
But stating it plainly like this would make the contradiction too obvious.
Though tbf corporate-speak and AI-slop are both insufferable in similar ways...
doesn't sound like a razor at all
that, they fully intend to 'pace'. hiring accenture is a good sign they need some justification theatre and a fall guy for this decision.
Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.
What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.
Simply make them something that derives a text response from its training data.
The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.
If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.
If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.
It does work out to be a similar cost per task though
If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate
For long running tasks it is. That's what made Deepseek so cheap.
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I've don't think Astra is a better model, but it's the first OpenAI one that seemed good enough for me. Definitely keen to try Opus 5.5 and see if this claim is real.
I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.
One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.
Giving moral lecture is different than reality i guess.
The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.
Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.
Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.
The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.
I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.
How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...
tired: AI startup attempting to publish a webpage
wired: a nonprofit founded in 1996
It's just a standard hero image + text for me, with no scrolling effects.
edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.
I agree that it's sort of stupid, not a fan.
For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.
Everyone that doesn't gets served some animated bullshit.
Yep, you're right. I tried on my phone and got the scroll through image.
> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!
God I hope so
https://github.com/AminBlg/SimpleEnglish
Nice. I was starting to think that Haiku got abandoned.
I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
Can you give more details here? This sounds intriguing.
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
This is news to me. Excited to try it out! Thanks.
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
This is where Chinese models are going to eat Anthropic's lunch.
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
All of this starts to feel more like a drug dealer selling their newest stuff.
In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.
And on the way I always have to check my tooling and need to adjust things to get max results.
Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.
Benchmarks often don't survive contact with reality.
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
Are the frontier labs even working on this problem?
I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.
This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.
Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Opus 5.5's output: https://html.non.io/annui-opus/
Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.
For comparison with other drops this week + current #1:
Astra: https://html.non.io/annui/
MiMo: https://html.non.io/annui-mimo/
Grok 4.7: https://html.non.io/Annui-grok/
Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.
The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.
Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.
For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.
Has oneshot all of the quite complex bugs / debugging tasks I gave to it which I know opus 5.0 would've struggled with
Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.
It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.
I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.
HN is a bubble that's mostly out of touch with what regular people use or care about.
In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.
Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.
Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.
[1]: https://github.com/p-e-w/heretic
IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).
Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.
They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?
Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too
It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.
It's frustrating that there isn't more transparency here.
Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM
Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.
I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.
> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.
> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.
> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.
It has the same annoying cadence and writing style with slightly less prominent claudisms.
* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.
* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]
* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]
and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.
I wonder if there are studies around this.
https://openai.com/index/where-the-goblins-came-from/
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs
"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"
and
"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."
and
"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."
I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.
If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.
Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?
How would this alleged difference (most likely bs) actually show up in reality?
GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.
>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??
Thank you.
It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.
Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.
Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.
---
I provided crapton of context for that one resume line. All the work I did, documentations for my justifications, etc.
I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).
---
As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.
Also adding verification for that Fable 5.1 output in the same sesssion.
Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.
Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?
Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.
[1] https://kizi.to/claude-talks-too-much/
Maybe this model can finally figure it out for them.
Not efficiency in writing, clearly.
Yay, yet another model I can't use for anything interesting, even with CVP.
Ants: It's a good model, sir!
We can't test it properly because it knows it's being tested.
Nice. I was starting to think Haiku was going to be abandoned.
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.
1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.
2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"
Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."
In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.
Not sure why but my guess is that this will be worse. Happy to be proven wrong.
I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.
Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.
"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.
The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...
It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....
I tried Opus 5 and Astra.
Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.
https://www.reddit.com/r/codex/comments/1wnggya/gpt_6_droppe...
Resets Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
What the hell does this mean? There are weekly "resets" anyways. And there will be 4 of them before Oct 22.
Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.
Holy shit! Its happening!
Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".
I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?
Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.
Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)
chinese models can't come soon enough
we're already getting enshittification
Less companies involved means less pressure to go fast.
Infomercial at its best.
No wonder we are hammered with ai announcements.
[1] - https://news.ycombinator.com/item?id=6372466