ZH version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
79% Positive
Analyzed from 4920 words in the discussion.
Trending Topics
#luna#model#more#models#cost#don#still#price#tasks#https

Discussion (191 Comments)Read Original on HackerNews
This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).
HN could be run as a BBS on 70's hardware. Instead of using a CPU with ~10 thousand transistors, you're likely using one with ~10 billion to do basically the same thing, and you don't think twice about it.
Famously, this is also a problem for human coders in sprint planning.
/s
I don't have the words.
I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.
The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.
https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.
Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen.
and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV.
It seems China is already able to do DUV a lot sooner than others expected.
Google is already working on a similar idea but more "flexible".
But if this is even at 400B size it's insanity those inference prices, maybe 10-20% margins, if it's higher I would like to know is it their own chips or maybe they have accurately sized the model to fit on exactly a B300?
Could be a lot of magical things we can only speculate, but from here there likely isn't another 60-70% margin, like I have heard people claim, I would definitely be willing to bet on that.
Could still be a healthy 10-30% margin. Especially with Terra.
[0] I am constantly surprised how much work pay-as-you-go with DeepSeek / MiMo will get done. I've barely crossed $2 each in a month of use (~200m tokens).
I feel perfectly content in using pay as you go pricing with deepseek. On the other hand, although Anthropic's models used to be my bread and butter for personal work, they are simply too expensive to reach for these days.
Assuming the efficiency gains are real, I feel like something has to give, maybe worse quality due to aggressive quantization/kv cache compression?
Although I'm sure there are some efficiency gains, the technology is too new and labs are scrambling to release too quickly to think that the low-hanging optimization fruit has been picked already.
If you use Codex it's different, the harness has a lot to do with it and there's definitely been changes including recently.
But yeah I do'nt want to know what Kimi 3 is pushing buttons inside Anthropic, OpenAI and Google.
Besides any floor: For every year the tokens get faster and cheaper, we will see new things like properly working AI factories which mimic expert teams. A lot more parallism as well.
Besides Nvidia Hardware is still sold out and super expensive. Not a single Nvidia consumer GPU got cheaper at all, Nvidia DGX Spark got more expensive too.
It will be swooped of the market the second it hits the market.
Luna is an extremely strong model.
Anthropic's big marketing push this year has been entirely focused on getting people to use Opus via a Claude Code subscription, to the point that Sonnet is almost viewed as the poor man's alternative, and from what I've seen, almost nobody uses it.
Actually, here's an interesting project for all the vibe coders looking for their next front page post: scrape a ton of commits from GitHub with Co-Authored-By: Claude and figure out what the percentage split between Opus/Fable/Sonnet is. I'm willing to bet it's less than 10% Sonnet.
This may be misleading, since I suspect many are using a blend through sub-agents. I tend to bias for Fable to orchestrate and Opus for implementation via sub-agents.
We are not purely rational creatures, thank God. Sometimes those "limiting factors" you listed -- stress, peer pressure, hormones -- are crucial elements of informing the problem solving process and arriving at a decision or a solution that actually works.
All an LLM can do is fulfill a prompt, no matter how misguided, backwards, or incomplete that prompt actually was.
"Go jump off a bridge." Hmm. Dying makes me stressed out. I'm not gonna do that.
Tokens are not normal software, because they have marginal cost, and I think people who are used to software economics really struggle with this. With token generation there really can be manufacturing cost efficiencies where one producer is just straight up better at serving product at a lower marginal cost.
With an 80% reduction in cost that becomes a ridiculous outlier in efficiency.
You mean they increased the price and then cut it back and now it is amazing?
Luna had a price hike vs mini (its previous replacement). The cut now just puts it back in that ball park.
Not that this isn't good news, but what's impressive?
I typically do lots of mini calls for research (100s of millions or something in that ball park). Newer models made that absolutely impossible, and the fact that the older ones are starting to get deprecated made me switch to e.g. deepseek for some of my runs. We'll see if I move back after this.
This is very likely priced below recovering the cost of the hardware but still above operating expenses.
I have no idea either way but one thing that detracts from these threads is folks claiming things as a fact without evidence.
https://www.youtube.com/watch?v=XDB5beon4DY&t=4m20s
> it doesn't feel like night-and-day.
I see what you did there. :)
Personally, I'm having surprisingly good results with DeepSeek 4 Pro at home, which is very good value for money: it's not as good as Claude / GPT 5.6 (I have Co-pilot license at work), but it's still really useful for code reviews, validating thoughts, and especially designing / writing unit tests for new (and old before refactoring) functionality.
And it's very cheap per task. (Flash is even cheaper, but I've had issues with that on more complex tasks where it starts forgetting things and arguing with itself "but wait, let me read the function again").
I was already a huge proponent of Luna for things like deep research. Being able to run 5x more for the same cost is simply bananas. We are already running 10 parallel agents for hypothesis generation. I cannot imagine 50. The statistics become much more interesting & powerful when you can run so many samples of the exact same prompt+model without breaking the bank.
Phase 1 - Run X copies of Luna in parallel over the user's prompt. The purpose is to generate a diverse set of hypotheses.
Phase 2 - Run Y copies of Terra in parallel to investigate the hypothesis results, with each receiving them in a randomized order.
Phase 3 - Run 1 copy of Sol over investigation reports.
The goal is to ensure that the agent covers more initial starting points before presenting a final conclusion. If you only run a single copy of Sol and it hooks onto something wrong, it might not recover.
Taking actions that mutate the environment is a different story. I think this is where you run into diminishing returns very quickly. You generally want one strong agent to act given the results of all the searching that was done. If the plan is clear, you don't need a genius model to execute it.
If the cost of serving GPT-5.6 just dropped by 20%, does that add up to literally billions of dollars in savings per month?
We know Anthropic spend $1.25 billion renting inference capacity from SpaceX (in two Colossus datacenters) from the SpaceX IPO, but we don't know how much of Anthropic's inference capacity that is (presumably a small fraction, since they were operating on top of AWS and other providers before the SpaceX deal.)
I've not seen any numbers that hint at OpenAI's per-month inference bill, but surely that has to be in the multiple billions of dollars as well.
So 20% is a really, really big deal.
> reduced inference cost by 20 percent saving company x billion dollars per month
I have no doubt that further work was required to enable this, but it's still very cool to be possible to say that.
I was still using GLM-5.2 in my personal projects, but this just made Luna a very easy choice.
I bet it must be better in some contexts and worse in others.
makes it by far the best choice for most workloads that do not need bleeding edge intelligence (reminder: luna can be comparable to opus 5!)
Haiku 4.5, on the other hand, is comparable to performance to Gemma4 31B (with working tool call formatting) in my experience, and Gemma4 strongly wins on vision and multimodal.
- Haiku: 30 points
- Luna Medium/High/Xhigh/Max: 38/46/49/51 points
That's a massive difference:
- 30 points is Gemma 4 31B territory
- 50 points is GLM-5.2 (744B) territory.
For our production app, OpenAI clearly is the best provider now. Their API is very reliable and has many nice features. The price-performance of the model lineup is incredible. We used open weights model via Fireworks for a long time (e.g. Kimi K2.5). Fireworks is a great provider but we still ran into issues here and there (Same with Anthropic and Google). OpenAI just works, is fast and in my view has a better price-performance ratio across almost all levels of intelligence.
Before I thought it was just an improved version or at least in the same class as gpt 5.4 mini but now it's being priced like a nano model!
I thought about it because Terra has similar pricing to 5.4 and Sol is similar to 5.5.
Luna was already my workhorse before, it performs very well on high/xhigh for most of the tasks, very happy about this drop.
I don't buy it.
There have been recent weeks where some of the mid-level models (Hy3, Laguna M.1) are free (true for parts of June and July, see Hy3 in Cyan) . Even then the total token usage appears to be reaching a steady-state.
https://openrouter.ai/rankings#top-models
^ the first graph is tokens per week across all models
I guess we just can only throw ideas at an LLM at a certain rate.
I still have ideas and now I can have an LLM vibe code what I want, but I'm not going to let an agent just run unattended for longer than a few minutes or a few bucks for hobby projects.
So maybe it is a matter of lowering the cost of an LLM so I can let it churn for hours at a cost of pennies... But I suspect demand for tokens is very price-elastic.
My company checks the models and pays for Opus through AWS.
You still send the WHOLE context of whatever you want to do to a random endpoint on the internet. If you want to write a good email, you give that context your email address, names, the reason for it etc.
Big companies don't randomly use some random api endpoint to do so.
Anthropics quarerly revenue is still growing very fast. I don't think we have seen even the real potenzial of it yet at all.
Not only are still a lot of countries missing which do not even use anthropic or any other frontier model yet but also all the agentic based solutions enterprise companies are currently building on mass (at least in my industry)
Your support bot.
Your research long running bot.
Your SEO Optimizer bot.
Your incident analyser bot.
Your personal assistent bot.
Whereas Google with Gemini 3.x, Anthropic with Fable etc are happy to just go for 'big model with dense params'
It's hard to guess from the outside of course but just this kind of talking points focus on GPU efficacy is what we see from OpenAI and Chinese open source labs more often than from Anthropic or Google Deepmind and this benchmark chart seems to concur
https://deepswe.datacurve.ai/ - (See the Agent Steps view)
Or is the output speed so much higher that it cancels out?
I don't see a lot of benchmarks that record actual time. But on AA, Sol on Low beats Luna on High for Time Per Task.
- lower input/output token pricing
- the cached token price is $0.0028/Million tokens, which is like 50-90% of tokens
presumably it's a much bigger model
Haiku was already in a ditch.
But this is coming straight for the jugular of a ton of models on openrouter.
https://files.catbox.moe/csxl32.png
(2 cents to run AA index, score 40)
Looks like OpenAI broke the pareto frontier on the trust-me-bro benchmarks!
(One has to wonder if they used any of the neat tricks from the DSV4 paper :)
With this new price change, Terra does look pretty Pareto’ed by Luna.
On agentic coding, pairing Sol Medium for architecting with Luna High for coding does kinda make sense. But beware that architecting can be very read-heavy, and Sol is a bit read-pricey compared to Terra.
Looks like the Chinese models are really making a dent. Having 3 different price categories with the "most affordable" one still costing more than GLM 5.2 never made sense.
> China: Household rates average around $0.08 / kWh (¥0.53/kWh).
vs
> US: Household rates average around $0.16 / kWh, though regional variation is massive—ranging from ~$0.10/kWh in low-cost states (like Washington or Louisiana) to $0.30–$0.45+/kWh in high-cost areas like California or Hawaii.
Estimated final electricity price for large industrial customers in energy-intensive industries:
USA 50 USD/MWh
China 68 USD/MWh
https://www.iea.org/reports/electricity-2026/prices
Edit: Yes, 80% minus is still milking. Because you empower these greedy mega-corporations. Just look at the RAM prices increase, then you see that the more money you give these hungry dragons, they more they will eat up. Don't get fooled by their "less cost now" advertisement.
Now, if cost is your concern, then that’s a problem in all of computing. Hence why I’m sending you short plain text messages using an iPhone with a many-core CPU and gigabytes of RAM.