DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
65% Positive
Analyzed from 2671 words in the discussion.
Trending Topics
#fable#opus#more#models#claude#better#sol#where#anthropic#different

Discussion (53 Comments)Read Original on HackerNews
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
But Fable security false positives and pricing just make it not worth compared to Sol IMO.
Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.
Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.
GLM 5.3 is a hair behind these two.
An anecdote but not an original one from the people I talk to.
I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a …
I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
Cheaper, more powerful AI will continue to expand the bubble. Projects will get more ambitious. Everyone will build out their own custom little software. Code diversity expands and requires even more AI.
They are already good enough at what they mechanically are.
You have to use the right harness, right verifiers (automatic where possible, human where not), etc much much more specific than a generic one like claude code or codex, and it will also be able to work within constraints and be the "proposer" of an imaginary optimisation problem and an excellent one at that. But you have to frame the task at hand in that manner or maybe even reorganise the task you do itself so it is more amenable to being framed that way. If you use it this way, it is _already_ massively economically useful. But it will take many years for it to actually be usable in that way, since you need DC capacity to come up first which is few years away and also well, massive organisations that have to integrate these will usually take many years to do so.
It is also useful albeit less so in cases like general SWE, where you still need a human in a loop for non-verifiable requirements, and also in other general usecases where information retrieval is too intractable and you need to carefully use LLMs as a component of the overall system.
I am not saying Fable or whatever the biggest models are are useless - they will certainly be useful for tasks at the frontier of the day - which is today complex exploits and open math problems, and well, tomorrow it could be something in biotech. But this is not what the entire bet is on at all - just automating day to day drudge at the tens of thousands of massive companies and governments we all know and love is more than enough. With the right training data (which _also_ is a bottleneck and takes time) you could even automate certain processes entirely. Sure, if we get a crazy medical innovation and end up saving trillions in healthcare great, but that's just a bonus.
None of this is to say that I think there is zero sketchy financial engineering going on
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use Fable still *feels* much better. I think benchmaxxing the new open weights models is ruining the value of benchmarks, if they ever had any. When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. And considering the point I'm poorly trying to make, that Opus doesn't feel like frontier anymore, this is probably the last month of my Claude subscription.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
I switched from Opus 4.7/4.8 to test Kimi K3 a few weeks back and the test hasn't finished; it's my daily driver now.
Given their general behavior and preference toward social engineering to scare the shit out of normal people...this seems fitting.
Nobody has complained and seems like for every use case we have Opus is more than powerful enough, especially with Opus 5
Most of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
No shit we're all paying 清冲 Flash to do the grunt work. Turns out though, paying 清冲 Flash a few more cents does exactly what Ivy Wasp Pro Mythical does. Crazy how that works.
I can literally open a new chat with just "Hello" and it gets bumped.
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.
https://openrouter.ai/stealth/ox-alpha