Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

79% Positive

Analyzed from 4920 words in the discussion.

Trending Topics

#luna#model#more#models#cost#don#still#price#tasks#https

Discussion (191 Comments)Read Original on HackerNews

GodelNumberingabout 2 hours ago
"Half the money I spend on advertising is wasted; the trouble is I don't know which half." -John Wanamaker

This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).

in_a_society42 minutes ago
I don't see why it should be all that difficult. All you have to do is first find a library that implements a decent solution to the halting problem and you're off to the races.
fractorialabout 2 hours ago
Cosmically apt username given the substance of this comment.
carimuraabout 1 hour ago
Exactly. I haven't reached the "let 1000 agents bloom" mode yet, so currently I'm spending real headspace managing agents doing work, and that work is all important, so why "settle" for sub-frontier models for that work? Maybe I'll get there for non-coding work.
pimeys13 minutes ago
If you have agents and users, you can run evals and see how far the models go. Luna is not greatest in tool calls, but if you define your problem well and the tools well, it is comparable to Gemini 4 Flash with much lower price tag.
satvikpendem8 minutes ago
Luna is good as an end user model for simple tasks like classification, but not as a coding model. Also do you mean Gemini 3.6 Flash? 4 doesn't exist, and Gemma 4 exists but doesn't have a Flash option.
bryanlarsenabout 1 hour ago
Highlighting https://news.ycombinator.com/item?id=49113236 in response.

HN could be run as a BBS on 70's hardware. Instead of using a CPU with ~10 thousand transistors, you're likely using one with ~10 billion to do basically the same thing, and you don't think twice about it.

throw2ih020about 1 hour ago
> separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).

Famously, this is also a problem for human coders in sprint planning.

londons_explore29 minutes ago
I get frustrated with a poor quality model leaving my codebase littered with wrong comments, which then later trip up smarter models.
odirootabout 1 hour ago
You just need a very strong frontier model to do triage of your tasks.

/s

wmf36 minutes ago
That's not necessarily a joke; the article proposes exactly that.
preommrabout 2 hours ago
> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less,

I don't have the words.

I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.

jpadkinsabout 1 hour ago
When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon and we will see another 10X improvement in price/performance frontier.

The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.

jrfloabout 1 hour ago
Burning the weights into silicon would be many orders of magnitude increase, not just 10x. It's kind of crazy that this hockey stick the AI hype bros talk about seems more and more every day like it might be real
jaggederestabout 1 hour ago
https://taalas.com/ has done it already for a wildly obsolete model. 14000 tokens per second.

https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.

Yopolo36 minutes ago
And don't underestimate how much money Google, Microsoft, Amazon and Meta still have to spend on this tech.

Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen.

and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV.

It seems China is already able to do DUV a lot sooner than others expected.

coffeebeqn22 minutes ago
What does that mean though? Like some kind of a ROM memory ?
re-thcabout 1 hour ago
> When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon

Google is already working on a similar idea but more "flexible".

captainblandabout 2 hours ago
To be fair we don't really know in terms of prices what's real and what's just investor subsidised attempts at market capture at this point. It could well be OpenAI's attempt to drown Anthropic while they've got the halo product if they feel they've got deeper pockets.
w29UiIm2Xzabout 2 hours ago
Enterprises implemented spending caps and inference providers are lowering prices. Seems they are jockeying for market share.
minrawsabout 1 hour ago
I wouldn't be surprised if they still had some margins since cheaper models are much harder to nail the accurate sizes off, and you still pay 2x for 1M context window.

But if this is even at 400B size it's insanity those inference prices, maybe 10-20% margins, if it's higher I would like to know is it their own chips or maybe they have accurately sized the model to fit on exactly a B300?

Could be a lot of magical things we can only speculate, but from here there likely isn't another 60-70% margin, like I have heard people claim, I would definitely be willing to bet on that.

Could still be a healthy 10-30% margin. Especially with Terra.

platinumradabout 2 hours ago
We can guess based on the decisions of other inference providers who serve these models.
handfuloflightabout 1 hour ago
Do you mean if other providers will cut their prices in turn?
foobar_______about 2 hours ago
Hard to believe numbers. I don't mean that as a critique, but literally I am so impressed. Even if the model is a few percent lower for performance but is 80+% cheaper than competitors and is a US company hosted on US based hyperscaler clouds this is kind of a no brainer. Hard for most businesses to justify otherwise.
rpdillonabout 1 hour ago
This is exactly the model that DeepSeek V4 Flash followed, and it's been insanely successful as a result, even though it's not frontier.
ignoramous26 minutes ago
DeepSeek v4 Pro & MiMo v2.5 Pro (Opus 4.6 quality models for code) are insanely cheap for agent-driven work due to their super low cached-input prices ($0.0036/mtok) [0]. For Luna, the cached-input price drop isn't disclosed in TFA, but the pricing page puts it at $0.02/mtok, & that's 5x more expensive.

[0] I am constantly surprised how much work pay-as-you-go with DeepSeek / MiMo will get done. I've barely crossed $2 each in a month of use (~200m tokens).

computerex9 minutes ago
Absolutely. Although DeepSeek started announcing "Peak valley" pricing which started making me nervous. I have spent $50 usd in July on deepseek and for that much spend I got SO MUCH mileage.

I feel perfectly content in using pay as you go pricing with deepseek. On the other hand, although Anthropic's models used to be my bread and butter for personal work, they are simply too expensive to reach for these days.

ismailmajabout 2 hours ago
it's 80% less cost, not 80% in efficiency gains, could be that Luna was overpriced to begin with, we don't have much info on the models themselves.

Assuming the efficiency gains are real, I feel like something has to give, maybe worse quality due to aggressive quantization/kv cache compression?

heisgoneabout 1 hour ago
Let's suppose each models was subsidized at 70%, so that we only pay 30% of the cost. They would loose much more money per token on the more powerful models. It's in their interest to encourage the use of the less expensive models. Let's say they increase Luna subsidies at 90%. They would still "save" relative to the use of the more expensive models.
mlinseyabout 1 hour ago
High-performing open weight models being released recently, and your customers looking into working with multiple providers as a result, are a great reason to drop prices on your non-frontier offerings.

Although I'm sure there are some efficiency gains, the technology is too new and labs are scrambling to release too quickly to think that the low-hanging optimization fruit has been picked already.

dannywabout 1 hour ago
Been using OpenAI models since ada/babbage/curie/davinci and at least from my own experience, their APIs feel the same.

If you use Codex it's different, the harness has a lot to do with it and there's definitely been changes including recently.

axusabout 1 hour ago
Something can be overpriced and still lose money.
Yopolo40 minutes ago
5-10% over months would still be quite crazy.

But yeah I do'nt want to know what Kimi 3 is pushing buttons inside Anthropic, OpenAI and Google.

Besides any floor: For every year the tokens get faster and cheaper, we will see new things like properly working AI factories which mimic expert teams. A lot more parallism as well.

827aabout 1 hour ago
Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.
cousinbryce35 minutes ago
In a data center that is power constrained but not space constrained they could build out new racks and flip the power from the old racks. Wonder if this will lead to moderately used server GPUs on the secondary market someday.
Yopolo31 minutes ago
I don't think they overengineered a DC like this.

Besides Nvidia Hardware is still sold out and super expensive. Not a single Nvidia consumer GPU got cheaper at all, Nvidia DGX Spark got more expensive too.

It will be swooped of the market the second it hits the market.

gentlewaterabout 2 hours ago
This is gonna put Sonnet 5 in a really awkward spot.
heaney-555about 1 hour ago
Luna is comparable to Haiku, not Sonnet.
827aabout 1 hour ago
Totally untrue. Luna and Sonnet 5 are very comparable: https://artificialanalysis.ai/#intelligence

Luna is an extremely strong model.

3836293648about 1 hour ago
Anthropic basically downgraded all their tiers when they released Mythos, no nah, Sonnet 5 is the successor to Haiku 4.x
baqabout 1 hour ago
I use sonnet as a smart grep and haiku never and that’s only when I have to use Anthropic at all
bakugoabout 2 hours ago
Sonnet and Haiku were already in an awkward spot, likely by design.

Anthropic's big marketing push this year has been entirely focused on getting people to use Opus via a Claude Code subscription, to the point that Sonnet is almost viewed as the poor man's alternative, and from what I've seen, almost nobody uses it.

Actually, here's an interesting project for all the vibe coders looking for their next front page post: scrape a ton of commits from GitHub with Co-Authored-By: Claude and figure out what the percentage split between Opus/Fable/Sonnet is. I'm willing to bet it's less than 10% Sonnet.

supern0vaabout 1 hour ago
>figure out what the percentage split between Opus/Fable/Sonnet is.

This may be misleading, since I suspect many are using a blend through sub-agents. I tend to bias for Fable to orchestrate and Opus for implementation via sub-agents.

StilesCrisis41 minutes ago
When I'm paying for it, Sonnet. When work is paying, Opus 5, then Fable if Opus gets confused.
petesergeantabout 1 hour ago
Opus 5 is not strong enough as the top-of-stack model, and feels idiotic after a week or two of heavy Fable usage, to the point where I'm paying for Usage Credits to keep using Fable rather than having to slum it with Opus.
arjunchint29 minutes ago
more like they were facing pressure from chinese models, and dropped prices and now their margins are squeezed
solarkraftabout 1 hour ago
They have no (other) equivalent to nano, so it makes sense that it’s much cheaper now. It may have been better, but it was also hell of a lot more expensive.
WarmWashabout 1 hour ago
Totally possible that humans aren't actually that intelligent.
ceroxylonabout 1 hour ago
As well as the existing intelligence being swayed by emotions, hormones, circadian rhythms, stress, peer pressure, propaganda, and survival instincts.
afry132 minutes ago
As if ALL OF THAT doesn't represent inherent and crucial components of "intelligence" itself.

We are not purely rational creatures, thank God. Sometimes those "limiting factors" you listed -- stress, peer pressure, hormones -- are crucial elements of informing the problem solving process and arriving at a decision or a solution that actually works.

All an LLM can do is fulfill a prompt, no matter how misguided, backwards, or incomplete that prompt actually was.

"Go jump off a bridge." Hmm. Dying makes me stressed out. I'm not gonna do that.

customguy37 minutes ago
That's a bit like saying a tail is swayed by a dog, as if it could exist without one, or would have anything to do if it did.
subw00f41 minutes ago
Why does it matter? This is completely based on data produced by humans.
camel-cdrabout 1 hour ago
this type of thing usually means you are the product
mediamanabout 1 hour ago
I don't see how this follows. The cost of nails has fallen by 95% over the last century. It's because the cost of manufacturing has fallen. Not because they are selling the information of nail consumers.

Tokens are not normal software, because they have marginal cost, and I think people who are used to software economics really struggle with this. With token generation there really can be manufacturing cost efficiencies where one producer is just straight up better at serving product at a lower marginal cost.

paytonjjonesabout 1 hour ago
According to https://deepswe.datacurve.ai/, Luna at Max at it's previous cost was comparable in both performance and cost to Sol at High.

With an 80% reduction in cost that becomes a ridiculous outlier in efficiency.

re-thcabout 1 hour ago
> Seeing spikes like this makes me question about where the floor really is.

You mean they increased the price and then cut it back and now it is amazing?

Luna had a price hike vs mini (its previous replacement). The cut now just puts it back in that ball park.

Not that this isn't good news, but what's impressive?

zzleeper37 minutes ago
Had to ctrl+f for someone saying this.

I typically do lots of mini calls for research (100s of millions or something in that ball park). Newer models made that absolutely impossible, and the fact that the older ones are starting to get deprecated made me switch to e.g. deepseek for some of my runs. We'll see if I move back after this.

visiondudeabout 1 hour ago
there is a ton of downward price pressure from Chinese open weight models
buckle8017about 2 hours ago
They over purchased hardware.

This is very likely priced below recovering the cost of the hardware but still above operating expenses.

infectoabout 2 hours ago
What evidence is there?

I have no idea either way but one thing that detracts from these threads is folks claiming things as a fact without evidence.

paxysabout 1 hour ago
That’s ridiculous. Every major AI lab is compute constrained. That’s exactly why nvidia is worth trillions today. If OpenAI had a single extra GPU they’d be using it to run another training cycle or experiment for their next model.
qntmfredabout 1 hour ago
sama literally just said they wish they had bought more. the price drops are almost certainly due to good old fashioned hardware innovation (wafer scale with cerebras) and optimizing hardware development based on model architecture and inference costs. other inference providers will try to do the same if they can.

https://www.youtube.com/watch?v=XDB5beon4DY&t=4m20s

pavpanchekhaabout 2 hours ago
Making Luna, which was already very cheap and extremely capable, 5x cheaper is crazy. I use Sol at work but Luna at home, and while there's definitely a difference, it doesn't feel like night-and-day. After a year of ever-increasing prices it suddenly feels (between this, Kimi K3, GLM 5.2) that prices are falling again.
jedbergabout 2 hours ago
> Sol vs Luna

> it doesn't feel like night-and-day.

I see what you did there. :)

deklesenabout 1 hour ago
Good observation! Kudos
pioneer37about 1 hour ago
Its just a matter of time at this point.These companies are working day and night to capture the market.
maxdoabout 2 hours ago
is kimi that cheap? it's a very expensive model
pixelesqueabout 1 hour ago
It's cheaper currently on many of the inference providers.

Personally, I'm having surprisingly good results with DeepSeek 4 Pro at home, which is very good value for money: it's not as good as Claude / GPT 5.6 (I have Co-pilot license at work), but it's still really useful for code reviews, validating thoughts, and especially designing / writing unit tests for new (and old before refactoring) functionality.

And it's very cheap per task. (Flash is even cheaper, but I've had issues with that on more complex tasks where it starts forgetting things and arguing with itself "but wait, let me read the function again").

subarctic17 minutes ago
I tried out deepseek v4 pro via a couple providers from openrouter, and it's always getting 429s. Are you running it on your own hardware?
dominotwabout 2 hours ago
depends on what you are doing. if you are doing verifiable tasks like fixing bugs then any model would do as long as you write the right verification.
bob1029about 2 hours ago
This feels like the dialup->broadband transition to me.

I was already a huge proponent of Luna for things like deep research. Being able to run 5x more for the same cost is simply bananas. We are already running 10 parallel agents for hypothesis generation. I cannot imagine 50. The statistics become much more interesting & powerful when you can run so many samples of the exact same prompt+model without breaking the bank.

jrfloabout 1 hour ago
Very interesting. Can you share more about your hypothesis/research pipeline? I have been using Sol for those types of task because I figured you'd need more reasoning for getting good ideas, but maybe quantity > quality at a certain point?
bob102942 minutes ago
Here is a rough approximation of the pipeline I use:

Phase 1 - Run X copies of Luna in parallel over the user's prompt. The purpose is to generate a diverse set of hypotheses.

Phase 2 - Run Y copies of Terra in parallel to investigate the hypothesis results, with each receiving them in a randomized order.

Phase 3 - Run 1 copy of Sol over investigation reports.

The goal is to ensure that the agent covers more initial starting points before presenting a final conclusion. If you only run a single copy of Sol and it hooks onto something wrong, it might not recover.

Imanariabout 1 hour ago
How do you run 'deep research'?
dannywabout 1 hour ago
Deep research is basically a LLM with web search, and a "work really hard" goal-orientated prompt, and some output formatting suggestions.
andaiabout 2 hours ago
Do you have a sense of which tasks benefit from more agents and which don't?
bob1029about 1 hour ago
Anything related to reading and interpreting the environment seems to always benefit from the addition of more agents to the search party, assuming you have some rational way to synthesize their results.

Taking actions that mutate the environment is a different story. I think this is where you run into diminishing returns very quickly. You generally want one strong agent to act given the results of all the searching that was done. If the plan is clear, you don't need a genius model to execute it.

handfuloflightabout 1 hour ago
I definitely think you want the genius model to synthesize everything that rolls up to them.
simonwabout 2 hours ago
> The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%.

If the cost of serving GPT-5.6 just dropped by 20%, does that add up to literally billions of dollars in savings per month?

We know Anthropic spend $1.25 billion renting inference capacity from SpaceX (in two Colossus datacenters) from the SpaceX IPO, but we don't know how much of Anthropic's inference capacity that is (presumably a small fraction, since they were operating on top of AWS and other providers before the SpaceX deal.)

I've not seen any numbers that hint at OpenAI's per-month inference bill, but surely that has to be in the multiple billions of dollars as well.

So 20% is a really, really big deal.

NitpickLawyerabout 2 hours ago
~2 years ago gemini2.5 helped write better kernes for itself and (only) reached 1% efficiency gains. Today we're at 20%.
dust4243 minutes ago
In 2 years from now we will be at 400%. https://xkcd.com/605/ Also, it is called kernels (you have nitpick in your username)
dominotwabout 2 hours ago
imagine writing that on your resume

> reduced inference cost by 20 percent saving company x billion dollars per month

paxysabout 2 hours ago
Where are you going to apply to with that resume that’s a step up from your current job though?
petesergeantabout 1 hour ago
The other place, but for more money
bpavukabout 2 hours ago
lots of places, actually. not everyone wants to be attached to the Silicon Valley culture, and that line alone will guarantee practically any workplace. that person is going to find out what work-life balance is :)
tekacsabout 2 hours ago
In this case, and I don't mean this critically, I guess it would technically be, "Instructed model to find efficiencies... reducing inference cost by 20% saving company x billion dollars per month."

I have no doubt that further work was required to enable this, but it's still very cool to be possible to say that.

andaiabout 2 hours ago
I think they meant that GPT-5.6-Sol can write that on its resume.
da_grift_shiftabout 2 hours ago
Does the model get the credit for its promo packet then? :^)
hirako2000about 2 hours ago
Contributed to. Can't be some IC who made a few nice PRs
quirinoabout 2 hours ago
I generally just check the Price/Performance graph on Openrouter: https://openrouter.ai/rankings#performance#benchmarks. Activate the "Show Pareto" toggle on the right.

I was still using GLM-5.2 in my personal projects, but this just made Luna a very easy choice.

qingcharles17 minutes ago
Hasn't OpenRouter had Luna and Terra on 50% off sale since they launched? I wonder what will happen to that.
hattimaTimabout 1 hour ago
The official doc says, Luna = Previous Nano models, kind of. Is it really good at coding?
PhilippGilleabout 1 hour ago
Depends on the reasoning effort, see https://deepswe.datacurve.ai (add Luna via model selection drop down, if it's not shown by default)
hattimaTim43 minutes ago
Thanks for the link!
paxysabout 1 hour ago
Smaller models are great if you are doing targeted changes in existing codebases. Don’t expect to use it for creating complex architecture from scratch or do major refactors. The larger the context, the greater the drop off will be.
quirinoabout 1 hour ago
According to the link I mentioned above it's roughly as good as GPT-5.4. Haven't tried it in practice yet.

I bet it must be better in some contexts and worse in others.

toshabout 2 hours ago
80% price cut for luna is a very aggressive pricing move

makes it by far the best choice for most workloads that do not need bleeding edge intelligence (reminder: luna can be comparable to opus 5!)

heaney-555about 1 hour ago
Luna is meant to compete with Haiku. What tasks are you seeing it equal Opus on?
dannyw22 minutes ago
You'd be surprised at what Luna can do, especially on xhigh or max. It's capable of working overnight, usually productively, just like Sol.

Haiku 4.5, on the other hand, is comparable to performance to Gemma4 31B (with working tool call formatting) in my experience, and Gemma4 strongly wins on vision and multimodal.

newtwilly43 minutes ago
According to the Artificial Analysis benchmark graph in the article, Luna can now outperform Sonnet 5 and Opus 5 low at ~4-10x less cost
euazOnabout 1 hour ago
Per Artificial Analysis:

- Haiku: 30 points

- Luna Medium/High/Xhigh/Max: 38/46/49/51 points

That's a massive difference:

- 30 points is Gemma 4 31B territory

- 50 points is GLM-5.2 (744B) territory.

toshabout 1 hour ago
luna is way better than haiku 4.5
nateb2022about 1 hour ago
I use Luna a lot (over 1T tokens since it came out) and I'd rank Luna (high/xhigh) on par with Sonnet 5, without hesitation.
__jl__about 1 hour ago
Didn't expect that. Luna pricing is crazy now. I don't think there is anything on the market that competes at this price-performance point.

For our production app, OpenAI clearly is the best provider now. Their API is very reliable and has many nice features. The price-performance of the model lineup is incredible. We used open weights model via Fireworks for a long time (e.g. Kimi K2.5). Fireworks is a great provider but we still ran into issues here and there (Same with Anthropic and Google). OpenAI just works, is fast and in my view has a better price-performance ratio across almost all levels of intelligence.

dannywabout 1 hour ago
OpenAI's APIs are extremely reliable for sure. I don't even remember when the last incident or downtime was.
amluto33 minutes ago
This doesn’t quite count as “API”, but OpenAI’s roll out of OAuth device code authentication was poor, to say the least.
Pesto16 minutes ago
I truly wonder what kind of model is luna now.

Before I thought it was just an improved version or at least in the same class as gpt 5.4 mini but now it's being priced like a nano model!

I thought about it because Terra has similar pricing to 5.4 and Sol is similar to 5.5.

Luna was already my workhorse before, it performs very well on high/xhigh for most of the tasks, very happy about this drop.

ninjahawk1about 1 hour ago
80% less for Luna is absolutely crazy, in my opinion we may reach a point in the next year where powerful models on the API could potentially be cheaper than subscriptions. Compute just keeps decreasing in price.
HDThoreaun31 minutes ago
API will never be cheaper than subs because theres a ton of value created for companies by locking people into subscriptions that tend to be sticky.
NortySpock39 minutes ago
> In a compute-constrained world where model demand is growing faster than capacity

I don't buy it.

There have been recent weeks where some of the mid-level models (Hy3, Laguna M.1) are free (true for parts of June and July, see Hy3 in Cyan) . Even then the total token usage appears to be reaching a steady-state.

https://openrouter.ai/rankings#top-models

^ the first graph is tokens per week across all models

I guess we just can only throw ideas at an LLM at a certain rate.

I still have ideas and now I can have an LLM vibe code what I want, but I'm not going to let an agent just run unattended for longer than a few minutes or a few bucks for hobby projects.

So maybe it is a matter of lowering the cost of an LLM so I can let it churn for hours at a cost of pennies... But I suspect demand for tokens is very price-elastic.

Yopolo25 minutes ago
These are free due to some different type of reasons like Nvidida sponsoring free tokens or the model companies.

My company checks the models and pays for Opus through AWS.

You still send the WHOLE context of whatever you want to do to a random endpoint on the internet. If you want to write a good email, you give that context your email address, names, the reason for it etc.

Big companies don't randomly use some random api endpoint to do so.

Anthropics quarerly revenue is still growing very fast. I don't think we have seen even the real potenzial of it yet at all.

Not only are still a lot of countries missing which do not even use anthropic or any other frontier model yet but also all the agentic based solutions enterprise companies are currently building on mass (at least in my industry)

Advertisement
Decabytes18 minutes ago
Has anyone ever done a comparison between the smaller models like Luna, against the previous GPT 5 frontier models? Have we gotten to the point where the small models are as good as the frontier models of the past, or is there still a way to go?
wronexabout 2 hours ago
What are your use case for these? I’m manly interested in coding where more capability is better - give me a 10x model at 10x the price and I’ll take it. A worse model at very low cost has no appeal to me. At-least not for coding. Translation maybe? OCR?
Yopolo25 minutes ago
Agentic layer.

Your support bot.

Your research long running bot.

Your SEO Optimizer bot.

Your incident analyser bot.

Your personal assistent bot.

wronex19 minutes ago
Game play bots (monsters, commanders maybe) would be really cool. But it needs long term support and probably local AI instead.
stri8tedabout 1 hour ago
Translation, moderation, classification, guardrails, etc..
firasdabout 2 hours ago
This is one of the things OpenAI has been focused on for an year or so that led to the doomed autoswitcher in ChatGPT .com (switching models based on estimated task complexity) that was quickly reverted

Whereas Google with Gemini 3.x, Anthropic with Fable etc are happy to just go for 'big model with dense params'

It's hard to guess from the outside of course but just this kind of talking points focus on GPU efficacy is what we see from OpenAI and Chinese open source labs more often than from Anthropic or Google Deepmind and this benchmark chart seems to concur

msejas30 minutes ago
Isn't OpenAI burning billions and have billions more spending commitments? If they managed to downsize so much the cost they should have kept the price the same and become profitable, really weird move, unsure what led to this.
otherme1238 minutes ago
Slower grow, or even shrink in usage? Right now, the promise of a future "everyone will use our models and pay whatever we ask" is what keeps $$ flowing towards OpenAI.
dgellow15 minutes ago
More than $650B due 2030. I don’t understand how it makes any sense that they reduce the price so much, unless they expect seriously such a massive saving and increase in demand from their latest improvements?
paxys15 minutes ago
Supply demand curve is a thing. Cutting price on something does not mean you are going to make less money.
dgellow13 minutes ago
But they still have to cover compute cost, and they already committed to more than $650B in infra expenses for 2035
Phlogi25 minutes ago
It's a clever strategic move: grab the market of cheap low end models within the product range. It's lower friction to switch a model than a provider.
andaiabout 1 hour ago
It says Luna is fastest, but doesn't it take way more steps to get the same job done?

https://deepswe.datacurve.ai/ - (See the Agent Steps view)

Or is the output speed so much higher that it cancels out?

I don't see a lot of benchmarks that record actual time. But on AA, Sol on Low beats Luna on High for Time Per Task.

arjunchint30 minutes ago
Deepseek Flash is still much cheaper:

- lower input/output token pricing

- the cached token price is $0.0028/Million tokens, which is like 50-90% of tokens

dgellow22 minutes ago
How is that economically possible? I’m so confused by those prices
anthonypasq14 minutes ago
how many times do you have to be metaphorically hit in the head with a brick before you realize inference margins at api pricing were 80%+
efficaxabout 1 hour ago
"too cheap to meter" and Luna is still more expensive than deepseek-v4 pro
jrflo44 minutes ago
It's marketing hyperbole, but Luna is more intelligent per dollar than deepseek-v4 pro. Cost means nothing without the associated capability
efficax39 minutes ago
is it? i don't know how we measure these things, but here's one measurement that says v4 pro is better than luna: https://artificialanalysis.ai/models/comparisons/gpt-5-6-lun...

presumably it's a much bigger model

energy12327 minutes ago
This has already been updated with the new prices?
xendoabout 1 hour ago
Does that also increase the token count in their subscription like ChatGPT GO?
Advertisement
incognito124about 1 hour ago
While I can't deny this is a huge technological result, and it's laudable they reduced the price because of it, 80% is really a lot. I can't help but wonder, is this because of the model's capabilities, or was the initial system just really sloppy? The public will probably never know the details
kingstnapabout 2 hours ago
Those prices on luna are killer.

Haiku was already in a ditch.

But this is coming straight for the jugular of a ton of models on openrouter.

gentlewaterabout 2 hours ago
This is awesome. I’ve recently set up my opencode to use 5.6 terra for my main agent, who delegates work to a 5.6 Luna coder agent. So far it seems to work well, and reduce costs a lot. With this price reduction, it will work a whole lot better. Perhaps I can get my github copilot quota to last the whole month now.
pehejeabout 2 hours ago
Might just resub. Will experiment with Luna next sessions. 5 h window is not working very well for me. But if I can drop down to Luna at 20-30 % left and comfortably ride out the wave then.. that might just work.
arcanemachinerabout 1 hour ago
They got rid of the 5hr quota, it's just weekly quotas now.
gck1about 1 hour ago
They're supposed to bring 5h today.
purpleideaabout 1 hour ago
I would pay significantly more to use these models if there was a legal contract that guaranteed they weren't ever terfing them and some way to prove that.
StilesCrisis30 minutes ago
What?
thehamkercat27 minutes ago
nerfing*
sosodevabout 2 hours ago
Looks like I might have a reason to use something other than Deepseek V4 Flash.
andaiabout 2 hours ago
I was curious so I photoshopped DSV4 Flash into the graph:

https://files.catbox.moe/csxl32.png

(2 cents to run AA index, score 40)

Looks like OpenAI broke the pareto frontier on the trust-me-bro benchmarks!

(One has to wonder if they used any of the neat tricks from the DSV4 paper :)

fractorialabout 2 hours ago
It would appear that rolling my own Anthropic-free harness / serving stack with a closed-weight carve out for Codex models is an absolute win.
jnakano89about 1 hour ago
Seems like they cut the tiers(GPT-5.6 Luna) where GLM and Kimi compete and still held margin for their frontier models
alvisabout 2 hours ago
Basically lunar at extra level can cover all use cases scenarios other than those requiring opus up. Goodbye sonnet and haiku
swingboyabout 2 hours ago
This is awesome. Luna is a pretty great model on xhigh.
Advertisement
goldsmith112about 2 hours ago
Not sure who would use Terra anymore. Pair Luna High/Xhigh with Sol Medium and that's your power stack
fritzoabout 2 hours ago
Sounds reasonable. Is there a good benchmark on which make this decision?
espadrineabout 1 hour ago
I maintain this meta-benchmark leaderboard: https://metabench.organisons.com/

With this new price change, Terra does look pretty Pareto’ed by Luna.

On agentic coding, pairing Sol Medium for architecting with Luna High for coding does kinda make sense. But beware that architecting can be very read-heavy, and Sol is a bit read-pricey compared to Terra.

andaiabout 1 hour ago
Sol as main agent, Luna for coding?
Aboutplants35 minutes ago
Your move, Anthropic
baalimagoabout 1 hour ago
We swapped an internal system from gpt-5-mini to gpt-5.6-luna and saw no benefit but 4x cost. Sufficed to say: we swapped back to gpt-5-mini.
gbnwlabout 1 hour ago
Experienced similar between 5.4-mini vs 5.6-luna in our own pipelines but after spending some time on prompt optimization and testing out various reasoning effort levels 5.6-luna was well worth it. Did you just replace model selection while keeping everything else in place or spend some time on evaling with newer prompts etc?
baalimagoabout 1 hour ago
No we kept prompts as is, just swapped model. The prompt is already quite optimized for the task. How would updating it possibly make a more intelligent model spend less tokens than a less intelligent model? Care to elaborate?
Tankenstein29 minutes ago
Most of the time when upgrading models we have needed to change prompts to get the same performance (let alone better performance). Usually, your prompt is overfit to the specific model doing the specific task. For example often your previous prompt is overspecifying and creating contradictions that a dumber model would just gloss over whereas a smarter model will try even harder to follow.
StilesCrisis8 minutes ago
So now presumably it'd at least be roughly equal cost, or maybe a little less?
hadlockabout 1 hour ago
Seems like they're working to destroy the local LLM argument. Right now Haiku is $1/$5 in/out. You can grind out $12,000 worth of haiku (or arguably, sonnet) class tokens in about 5 months on a Blackwell RTX 6000 96GB especially if using concurrency. BUT, but, if you use a g6e.xlarge on aws it's now more expensive than buying tokens from OpenAI @ $0.20/$1.20. It also destroys "the Mac Mini argument", pushing the ROI to ~4 years.
jrflo42 minutes ago
The local LLM argument never really held water tbh. You can get surprisingly good performance for lightweight tasks locally, but you're just fighting economies of scale if you're going trying to beat a datacenter on cost.
bakugoabout 2 hours ago
> GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less

Looks like the Chinese models are really making a dent. Having 3 different price categories with the "most affordable" one still costing more than GLM 5.2 never made sense.

preommrabout 2 hours ago
I thought the chinese models were cheaper per token, but about the same or more expensive on tasks because they used more tokens for reasoning. Cutting even further, seems like a really big leap.
measurablefuncabout 2 hours ago
It all comes back to electricity cost. China has cheaper electricity so as long as China keeps pace there is no way for American companies to undercut them. Each boolean operation in China is cheaper than the one in America.

> China: Household rates average around $0.08 / kWh (¥0.53/kWh).

vs

> US: Household rates average around $0.16 / kWh, though regional variation is massive—ranging from ~$0.10/kWh in low-cost states (like Washington or Louisiana) to $0.30–$0.45+/kWh in high-cost areas like California or Hawaii.

cbg0about 1 hour ago
This doesn't seem correct.

Estimated final electricity price for large industrial customers in energy-intensive industries:

USA 50 USD/MWh

China 68 USD/MWh

https://www.iea.org/reports/electricity-2026/prices

tokai36 minutes ago
I don't know, non of the chinese models I use are served from China. And they are still cheap.
shevy-java28 minutes ago
The milking games have started. The billionaires want their money back.

Edit: Yes, 80% minus is still milking. Because you empower these greedy mega-corporations. Just look at the RAM prices increase, then you see that the more money you give these hungry dragons, they more they will eat up. Don't get fooled by their "less cost now" advertisement.

measurablefuncabout 2 hours ago
Model segmentation & distillation like this that asks the consumers to pick exactly which version of the algorithm will solve their problem is evidence for lack of intelligence instead of its presence.
beeringabout 2 hours ago
You really really don’t need to pick. Just use Sol on high. That’s my daily driver and I don’t touch the model picker at all.

Now, if cost is your concern, then that’s a problem in all of computing. Hence why I’m sending you short plain text messages using an iPhone with a many-core CPU and gigabytes of RAM.

dominotwabout 2 hours ago
it is really hard to know upfront if you have fuzzy task. sometimes i would choose a cheaper model and it will spin and spin with bad outputs ending up costing more had i chosen a more capable model.
cute_boiabout 2 hours ago
there is mixture of experts which is also another routing. So, simple change in prompt can be a big difference.
sidcoolabout 2 hours ago
They don't mention Grok at all.
andybakabout 2 hours ago
Don Draper in the elevator meme?
hirako2000about 2 hours ago
Of course. All comparison is with what makes them look good.
paxysabout 2 hours ago
They also don’t mention a hundred other models.
qingcharles19 minutes ago
Musk announced Grok 4.6 coming next week, no idea what changes that brings or how it compares to the current 4.5.
wilgabout 2 hours ago
What would they say about Grok?