Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
78% Positive
Analyzed from 3197 words in the discussion.
Trending Topics
#models#model#don#more#run#local#llms#better#llm#slms
Discussion Sentiment
Analyzed from 3197 words in the discussion.
Trending Topics
Discussion (70 Comments)Read Original on HackerNews
Read the paper the author talks about yourself instead (https://arxiv.org/abs/2511.07885), and also, contrary to what the author says in the article, do not do investments based on single papers made from academic studies, regardless of how much money this guy tells you you can make.
The compute capacity, however, is here to stay. There will be a seemingly endless line of people queued up to buy it at pennies on the dollar if this whole thing blows up financially. And they will use it!
Just out of curiosity: Use it for what?
I'd disagree here. I see two avenues for an efficiency multiple, albeit a single-digit multiple:
* Client aggregation allows a hyperscaler to average out demand spikes from uncorrelated clients, reducing the peak:average demand ratio and allowing better budgeting of compute.
* Dynamic batching allows typical requests to run in batches of more-than-1 and/or overlap, offering better internal compute utilization ratios (e.g. interleaving output and input streams). The small limit of on-device LLMs will run with batch sizes of one with strong memory bandwidth bottlenecks.
For an example of these factors in action, see the API cost differential between batch, standard, and 'fast' processing. OpenAI prices these tiers at a 1:2:4 ratio.
Maybe if you combine an SLM with a database (as a tool) then it could work, but someone should first prove that.
1. Training a large model with lots of information, then stripping the "useless" information from that model to obtain a small model => nobody has shown this.
2. Training a small model, letting it use a database tool so it scores the same as a large model without database => nobody has shown this.
But seemingly models good at programming for example, would get worse at programming if you removed everything not-programming. Train a model solely on syntax, and it'll be worse than a general purpose LLM on syntax, in general at least.
There may be a great deal of advantage to be able to run a small language model on something I have in my hand, disconnected.
Although mainframes have their use, the pendulum of centralized to distributed has gone back and forth and there are benefits to be gleaned from either model, sometimes at the same time.
Also, the average consumer is not going to be running a local model until they are built into the hardware they already buy and when they are, who is supplying the weights? They’ll likely be shipped as an ASIC (or MSIC) at that point anyways. Those will use a licensed model from the current leaders. The whole argument sounds like saying that cloud services shouldn’t be profitable because everyone has a computer at home or to meme “we have AI at home”.
Apple or Nvidia, presumably.
Says who? In the world of video decoding, H.265 is losing out to AV1 largely because it's not superior enough to H.264 to justify the expensive and complicated licensing.
Do you really think chipmakers are going to pay a 25c per unit tax to get a model that is 15% better?
1. The hyperscalers are in a positive reinforcement loop. Despite any suggestions to the contrary they keep getting bigger. And can, er, “influence” government policy/officials and anything else needed to keep it that way.
2. The frontier labs and their investors. Another self-fulfilling reinforcement loop. Witness the circular gymnastics among OAI/Anthropic, Microsoft/Amazon and Nvidia
3. Data. No-one believes that Zuckerberg and co are going to say “great, we can just run the models on devices we don’t own and stop the surveillance economy because, y’know, privacy matters and we really care about mental health”.
And then there’s data centre locations and “yeah but jobs” even though your power bills are going up, and “why run your own data centre Mrs CTO, let us do it for you and save all that capex and those pesky employees you need to do it”.
Don’t get me wrong: I’m rooting for local, open weight/source models. But “hey look they benchmark well” is an unhelpfully narrow basis to forecast the demise of central hyperscaler hosting.
How many developers here don't see a difference between the latest LLMs and SLMs they can run on their own computer? I tried running a smaller model locally, and it's not usable for me.
I know people like to "predict" things, so that if they happen they can then say "I am a visionary, I predicted it" and start their blog posts with "as I predicted long ago (because I am a visionary), ...".
> The research report estimates that the addressable market in the US for SLMs has grown to about $10tn or one-third of the entire US GDP of $30tn. There isn’t much left for LLMs to thrive in, and every year, their advantage over SLMs is shrinking.
I stopped counting the number of times "estimates" said that a market would absolutely explode, and it absolutely didn't. Those are in the business of being a broken clock.
If something better comes, it will be better. Sure. And we would like to have something better, because it would be better.
If you have the hardware, a MacBook Pro for Qwen 3.6 35B A3B and Gemma 4 26B A4B for example, they are absolutely usable, both in terms of speed and quality. Anecdotally, I can use Qwen for day-to-day coding tasks in TS and Go, without hickups.
With local models, this iteration cycle takes maybe 30 minutes for a single fix or feature, rather than 10 minutes with GPT+Codex, as there is so many corrections and iterations needed, although I will say that the speed I'm able to get locally makes it more fun that any of the remote models.
This is becoming increasingly important to me. Super smart max reasoning frontier is fine if I leave it running overnight on some prepared set of clearly defined tasks, but when I want to work with the LLM, throughput really matters, and I'll go with a dumber model to get there.
At some point though, it's fast enough and any speed gains beyond that just makes me the bottleneck.
I also am seeing the smaller models gaining big strides lately, closing the gap on frontier models (still a decent sized gap though). I don't even run the small models like Qwen 3.8 27B locally. I just try them out in the cloud to see how they are progressing, and I'm definitely able to be productive.
The lack of logic and risk management on this statement, is so strong, I hope humans are all quickly substituted by LLMs. Lets just do it and be done with it...
Presumably not what you intended but this phrase immediately takes me to:
https://www.youtube.com/watch?v=dJFR7xbOIuw&t=42s
Why don't you go talk to your LLM instead of commenting here, then?
Probably a skill issue on your part.
If this is correct I see a future where the hyperscalers are funded by the businesses integrating siloed SLMs in their software.
Also the defence/intelligence industry will always want to keep an edge so don't be surprised if they stick around and we see favourable regulations for them similarly to how the government turns a blind eye to social media platforms because they increase the footprint of mass surveillance.
I wouldn't be surprised if the hyperscalers became software auditors and any piece of critical software was required to have a regulated security audit before it could enter production. Selling the poison and the cure is a great business model.
You can imagine a reasoning model as a huge set of rules that generate the next statement from previous statements (written in context). In that sense, a reasoning model can be compared to a logical theory - you have certain deduction rules which can generate new judgments.
Often, logical theories are structured that the rules are remade into axioms, and the deduction rule is only modus ponens (which corresponds to function application and is a building block of program execution).
In the case of an LLM, the set of rules (or axioms) they have in the theory is quite large, but most likely semantically unsound (with respect to their their own representation of truth) - that's why LLM's make mistakes.
It would be desirable to break the logical theory represented by LLM into a smaller set of axioms, which would:
a) remove rules easily deductible from the smaller core of axioms (for example, LLM doesn't need to remember "Socrates is mortal", as it can derive it from "Socrates is a man" and "all men are mortal")
b) remove rules that have low value (facts that aren't used often or have weak validity) which cause ruleset to become unsound
I suspect that's what SLM distillation is doing, to some extent.
The question is, how far this process can go? I personally believe there is a useful logic for commonsense reasoning that has less than thousand rules (still several orders more than your typical mathematical logic, but orders less than SLMs). These axioms do not contain much facts about the world, but that could be added.
So I believe there is a sweet spot (deductive core, encyclopedic shell) which we have not yet found (it's a little bit more formal language than natural language) but is very efficient for general reasoning.
Then of course there is the economics of it. Do people prefer to spend $5000 upfront to get things done 5x slower, or would they rather pay $20 a month for that?
If you think of for e.g. some proprietary piece of software that wants to embed an LLM they've fine tuned or trained, they will want to make back some of their research cost right. So they are not going to want to put this on-device even if the hardware is there, unless there's some way of locking it down. I suspect we'll need on-hardware validation/verification and a way of preventing extraction of weights for this move to happen for many use cases.
An implication is that successful research in "I don't know" detection could destroy hundreds of billions in shareholder value.
First time I hear that...not really true.
"Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors" - https://arxiv.org/abs/2607.00447
"Calibrated Language Models Must Hallucinate" - https://arxiv.org/abs/2311.14648
"TruthfulQA: Measuring How Models Mimic Human Falsehoods" - https://arxiv.org/abs/2109.07958
What might happen is that a chunk of the market, whatever its size will be, will end up going to SLMs run on iphones or Macbooks, and eat some of the revenues from LLMs, because not everyone needs the most powerful LLM all the time.
Imagine current frontier models at 20k tokens/second.
> they provide a better or at least as good an answer as LLMs in 62.5% of the cases.
Are we going to scrap hospitals because a vet could do the job 62.5% of the time?
The economics also point away from everyone buying a big RAM Mac that sits idle 99% of the time. SLM and own hardware sounds efficient and “free” but it is nothing of the sort when you factor everything in (and forfeit the sharing efficiencies of API)
SLMs are great esp for task specific fine tunes but this take isn’t it
What if they get sufficiently powered and watered industrial warehouses close to where the successful people live?
So maybe in 10yr I'll be able to run a SLM on a 5yo laptop and not have Google or whoever hoover up everything.
So what I’m reading here is “LLMs have a significant advantage” in the most critical areas that have practically infinite demand for more intelligence.
This means that the Frontier labs are under immense pressure to maintain that lead, and could end up in serious trouble if they stumble at all.
The other thing id point out is that a lot of us who are token-sensitive do things like build plans using expensive, smart models, and then execute those plans using cheaper dumber models.
Then there's the fact that we are still in the age of heavily subsidized Frontier subscriptions + tokenmaxxing initiatives from megacorps. Neither of which are sustainable, and will drive more usage to smaller open models once they end.
It is not the whole story, and knowledge is very lacking, but it has gotten a lot of attention. That model together with DeepSeek V4 Flash are the highlights of this summer on the open/local models side.
There is a likely US scenario and a rest of the world scenario. It will be interesting to see if China acts on the overextension of the US Military in the Middle East. Taiwan will be a big play for both and crucial to the hyperscalers.
But since the US is dabbling in piracy again and telling people what they can do and not do with their shit, it’s not too far-fetched that everyone that is not a global superpower is at risk of getting bombed to smithereens if they are a danger to US AI supremacy.
This is such a crazy timeline, predicting even like a single year ahead feels like looking into a medieval glass ball.
But we are humans, I am confident we will find a way to fuck this up royally for everyone. Brace for impact.