HI version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
88% Positive
Analyzed from 1637 words in the discussion.
Trending Topics
#glm#flash#models#deepseek#model#price#cost#more#agents#same

Discussion (42 Comments)Read Original on HackerNews
- Because it's a heavy reasoner, it sits near Gemini 3.7 Flash on the Pareto front (not as cheap as the price suggests in practice).
- Closer than expected to the top open weights models (GLM 5.3 and Kimi K3) in agentic coding, at lower cost.
- Chinese models have always been strong iterators in an agentic harness. This model is no different, reaching an average percentile ~20% higher when given a harness vs a one-shot solution. That one-shot fluid intelligence is what makes a model feel smart, though, and typically results in fewer attempts/tokens to solve a problem, and American frontier models are still far ahead in that department.
The new architecture is interesting. It puts pricing between their old Flash and Pro lineups, suggesting they might be abandoning their super-cheap flash models (which weren't that fast due to heavy reasoning) and their pro models (which sort of flopped and weren't consistently better than their flash models, despite the size/cost) and shipping a strong intermediate that competes with the Gemini Flash series.
Data at https://gertlabs.com/rankings
I just haven't found them to be very good? I've had a ton more success with the GLM models (since 5.2 anyway). Maybe I'm just holding it wrong, DS models seem to get stuck in loops or tell me nonsense. GLM feels like budget Claude.
It's really crazy. We price per token our customers. If Kimi was about 30% of the price of Opus 5 for the same quality, DeepSeek is 1/10th of a price of Kimi K3. We've come down in price so much that I seriously cannot recommend other models before they reduce their pricing.
And what is really interesting is its programming ability. As I've said in my previous comments, I use agents a lot in my work. Since early Opus days until now I have 7-8 agents working in parallel for different tasks. Rust, design, GEPA, evals, analysis. For a long time Kimi K3 was the best model for this work, and before that GPT 5.5. But I still can't really believe how well DeepSeek works here. I really try to find faults from it, trying to see that it must be doing sloppy work and be worse than the others. But it does not. It finishes every task I give to it. And the cost per task is under a dollar, usually 15-30 cents.
In comparison the same task with Kimi would be 3-15 dollars; sometimes closing to 100. And before that with GPT 5.5 a 800 dollar task was not uncommon if I spent days evaluating models.
Now it's less than a dollar.
For me if the other providers will not drop their prices dramatically in the coming weeks I see no reason to use them. Even with a 200 dollar subscription, paying per token for DeepSeek is better value.
My harness: https://omp.sh/
I have never had looping issues with DeepSeek models. Which provider/serving framework and harness are you using?
I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.
Perhaps DS works better where targets have low-hanging fruits than can be found fast?
Given how fast and cheap DS is, it's just an ideal model with enough "IQ" to let it loose. Another thing they left out of the article, DS becomes really good if you provide custom tools for the task, on it's own it's mediocre.
Inference providers can credibly promise to not train on your data if they are in a position to get sued.
But, well, number of subagents doesn't make a difference if model is dumb (GPT 5.4 High, in May,in Chat mode outperformed what I see with DS4.1-F).
That being said, pricing model makes a huge difference for "find at least one" tasks: with API/PAYG if you have a chance to save 90%, you go for it, whereas with subscriptions it is optimal to burn all your remaining allowance right before reset
Therefore you use for offensive cybersecurity tasks because Daybreak Red/Mythos is pure unobtainium for us mere plebians.
DB Blue thankfully exists, but I suspect you risk a ban if you use it with codebases you neither own nor use
Tl;dr because it's the only model at the level of 5.4~5.6 that doesn't refuse tasks nor risk your oai account getting banned
For plain RE tasks Sol or Astra should work just fine (I think)
I love DS flash, it is an amazing workhorse to implement plans created by more robust models (such as GLM). But a more fair comparison would be of DS Flash with GLM Flash.
I have good results with DS4.1 flash because I can iterate faster. I either provide it with correction, or it discovers its failures via the harness. And seems to respond well to empirical evidence rather than go in circles.
So it might need some prodding, but it's likely in this case it was able to brute force after several runs and collecting some evidence.
The 2 trillion dollar ROI on anthropic alone?
Good luck with that.
> . The accepted runs cost $4.65. Failed attempts and replacement runs increased the complete cost to $5.14.
Come on. Lecturing about hubris when your message is damn the consequences, full speed ahead?
If it's a race to build the torment nexus, or a race with a nonzero chance of building the torment nexus by accident, I don't care about winning.
They said that... where?
By dismissing concerns and focusing exclusively competition and "winning the race."
“If we ban CFCs now the Chinese will win!”
“If we ban chemical weapons, nuclear weapons, etc etc our enemies will triumph! They won’t stop!”
“If we switch to biodegradeable plastic then our rivals will have an advantage.”
“If we dont externalize the costs to our population, then they will, and then will win!”
I think workflows can do the job agents do, 20x cheaper and more predictably and safely. They can completely displace agents, just as HFCs displaced CFCs and then we were able to ban CFCs and phase them out through international COOPERATION. The language of COOPERATION is what saves us vs COMPETITION is all about cutting corners and externalizing costs. Google the Montreal Protocol, Geneva Conventions, Nuclear Non Proliferation Treaty, Unleaded Gasoline etc etc.
Agents have got to be marginalized. They are just popular because the labs need to make a ton of money for their investors and recoup their massive spending on training models.
You can't compare banning football to banning genetic experiments and say "they are both bans and therefore directly comparable"
We're talking about replacing most uses of Agents with uses of Workflows. That's the key first step, without which the industry will cry "don't regulate us... do you want China to WIN???"