Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
74% Positive
Analyzed from 9939 words in the discussion.
Trending Topics
#sol#model#models#astra#opus#gpt#more#openai#better#don
Discussion Sentiment
Analyzed from 9939 words in the discussion.
Trending Topics
Discussion (401 Comments)Read Original on HackerNews
Subagents are like trading derivatives. You can lose as much as you want.
When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.
Who cares if your car can go 200mph if all you need is 60. If my requirement is 60mph, I want a faster 0-60, not a higher top speed.
Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.
I use it as main Hermes model that orchestrates codex/droid harnesses with subscriptions for heavy dev work
I do have ChatGPT as main assistant that sets direction and delegation of projects to Hermes
At my increasing usage, kind of 200 usd subscriptions makes sense and max out on Luna max
I get the same UX on every platform, works perfectly on very low bandwith environments such as in a cabin, in the subway or in the middle of nowhere.
I tried using other harness such as Pi and opencode but I did not like them. If Claude Code gets weird I can swap in an instant.
You just need to follow this guide and disable artifacts in Claude Code's config: https://api-docs.deepseek.com/quick_start/agent_integrations...
Not sure that's why they did it. But that was my experience.
The smoking gun is how much slower than Sol 6 this is. It's not a retrain.
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
However it's less willing to obey your instruction so it's less usable for general runtine flows.
(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)
1 - https://bench.killswitch-lang.org
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
do you think it will be exponential forever?
What a time to be alive.
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
- it’s correct there isn’t much fresh data anymore
- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive
- it’s correct the finances don’t make sense
But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
Edit: removed a comment that was uncharitable and rude, for which I apologize.
We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.
People were talking about plateau for years already.
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
It's not even anything controversial..
I remember when bandwidth was super expensive and now it’s dirt cheap.
Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975
it's also why there have been so many calls for regulation and slowdowns.
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
I just added an agent / coding agent into an email app, and doing it through `codex` and its Codex App Server couldn't have been easier, and the results are very compelling.
The open source harness, API around it, and friendliness for connecting a subscription puts Claude to shame right now.
A few more thoughts here https://housecat.com/blog/introducing-housecat-agent
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...
Yikes
Our VC-backed subscription days are numbered
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
Tibo said that the existing $200 subscriptions keep the 20x factor for a while.
Ultrafast would have been nice with the temporary "Pro 400" plan.
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.
That they are expensive and climbing doesn't negate my point if the cost of the subscription over how long you plan to keep it is equally or more expensive than the GPUs. You can put together dual 5060 Ti or 5070 Ti systems to run local LLMs too. You don't need to splurge on a 5090. That's a bad option at this point.
I can justify $200/mo but more than double is not appealing to me.
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
You have to do a lot of things in parallel.
Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.
I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.
Whether that survives contact with Chinese open weight models is hard to say.
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
The only way is for prices to go up. Way up.
See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring for some.
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
A number of others have done game/3d video benchmarks but this guy is probably the most prolific.
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.
https://artificialanalysis.ai/?models=gpt-5-6-luna-low%2Ccla...
According to this, at Max it's better and cheaper than 5.5 Medium, but worse than 5.5 High. At Medium, it's better and cheaper than 5.5 Low.
Excited to tryout Decisions API as well.
Astra is a pretty impressive model. Excited to try this.
Opus 5.5 was a gut punch and my impression is OpenAI is still reeling.
The best thing is that we benefit from these constant back and forth gut punches :)
They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
Back fired because of opus 5.5.
So now we get the real sol-6 as sol-6.1, and OpenAI will eat the cost to stay competitive.
This could be invalidated if sol-6.1 is the same speed as sol-6.
However, that doesn't say much. You can just run a smaller model at a larger batch size to get higher throughput but lower interactivity.
The last time a model announcement felt like a leap in capability beyond other things out there was Fable - which was promptly taken away. Sol and recently Opus 5.5 were strong because they approach that capability with a lot more efficiency and don't blabber incoherently (looking at you Opus 5.1).
Deepseek is a workhorse for those who prefer open and API usage. Other than that the model announcements all just seem like a blur and quite interchangeable but I wonder if that's just me tuning out or do others feel the same way?
My experience with agentic coding on projects I care about (because my responsibility in my firm is to care about these things, at least for now) has not changed a lot in the past few months, and I have kept up with every single model update / experimented with harness a great deal.
Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr...
Interesting that High got the render order correct, with the back leg behind the bike, while xhigh and max have both legs on the same side of the bicycle. Astra only got this right on Max.
Either way, a little ironic…
These moves all make sense when you take into account the enterprise market.
https://news.ycombinator.com/item?id=49889873
1 - https://bench.killswitch-lang.org/
The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable/Claude just cannot get/fix even when you point it.
This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.
Will see if this remedies things.
https://amphetamem.es/meme?id=the-simpsons_06_12_71&text=We%...
Fuck altruism, ammi right? lets make money, gobs of it by screwing the middle users as much as we can to push them into just two tiers: Ones that use it for recreation and others that pay through their noses.
If OpenAI cuts alternative harness support it will be a weird day trying to figure out what to do next, it's been so clearly the best bang for your buck (imo) for a while. maybe id finally have to give smaller models a try.
anything to avoid using the dogwater codex & claude code tuis.
anyways this seems like a nice cost improvement over GPT 6 Sol and I expect this will be my new daily driver.
not saying this is the case here but it does feel a bit like wine tasting sometimes, everyone claims to be an expert that can taste a few tokens and tell you exactly what region and vineyard its from.
One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.
In other news:
> In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.
Huge misstep releasing it.
Guess not?
{"type":"item.completed","item":{"id":"item_0","type":"error","message":"Model metadata for `gpt-6.1-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues."}}
I can handle issues much better if they are predictable even if the model makes mistakes — much more frustrating when the model is erratic
I find codex wanders off road more often and fails to see the “bigger picture” (as much as LLMs can see the bigger picture at least)
And tbh when it was first released Astral felt even worse
I’m being forced to use it right now and at the end of the day I’m making do so it’s fine, but Claude makes for a smoother experience
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
Edit: for context, just Steam alone has ~200million monthly active users.
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
How many DCs are devoted solely to gaming?
An entire planet. Just Steam alone has one or two hundres million monthly active users.
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?
> I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Just Steam has 200 million monthly active users. Add Steam, PS, Xbox, and whole other devices having gpus and I'm pretty sure you at least 10x the energy consumption of all ai companies.
Yeah? Show me the big movements against computer gaming.
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.
When do we get the next jump in capability? When is 6.1 Astra released?
After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.
The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.
Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.
Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.