Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

86% Positive

Analyzed from 1155 words in the discussion.

Trending Topics

#models#model#local#frontier#don#llms#better#advantage#slms#using

Discussion (12 Comments)Read Original on HackerNews

Animats•about 2 hours ago
A remaining advantage of large language models is that as they get larger, they tend to hallucinate less, simply because the odds of the training set containing a desired answer improve with size. If a solid "I don't know" detector is developed for inference, then you can try a small language model first.

An implication is that successful research in "I don't know" detection could destroy hundreds of billions in shareholder value.

root-parent•about 2 hours ago
>> A remaining advantage of large language models is that as they get larger, they tend to hallucinate less

First time I hear that...not really true.

"Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors" - https://arxiv.org/abs/2607.00447

"Calibrated Language Models Must Hallucinate" - https://arxiv.org/abs/2311.14648

"TruthfulQA: Measuring How Models Mimic Human Falsehoods" - https://arxiv.org/abs/2109.07958

embedding-shape•about 2 hours ago
Another "cool but we don't know how yet" thing would be a "confidence interval" so we know how much to trust LLM responses. Or while we're fantasizing, they could just know everything all the time regardless of training data. The "if a solid" part is easy to imagine, hard to implement :)
CTDOCodebases•about 2 hours ago
Haven't the SLMs been distilled using the LLMs?

If this is correct I see a future where the hyperscalers are funded by the businesses integrating siloed SLMs in their software.

Also the defence/intelligence industry will always want to keep an edge so don't be surprised if they stick around and we see favourable regulations for them similarly to how the government turns a blind eye to social media platforms because they increase the footprint of mass surveillance.

I wouldn't be surprised if the hyperscalers became software auditors and any piece of critical software was required to have a regulated security audit before it could enter production. Selling the poison and the cure is a great business model.

palata•about 3 hours ago
"If", sure.

How many developers here don't see a difference between the latest LLMs and SLMs they can run on their own computer? I tried running a smaller model locally, and it's not usable for me.

I know people like to "predict" things, so that if they happen they can then say "I am a visionary, I predicted it" and start their blog posts with "as I predicted long ago (because I am a visionary), ...".

> The research report estimates that the addressable market in the US for SLMs has grown to about $10tn or one-third of the entire US GDP of $30tn. There isn’t much left for LLMs to thrive in, and every year, their advantage over SLMs is shrinking.

I stopped counting the number of times "estimates" said that a market would absolutely explode, and it absolutely didn't. Those are in the business of being a broken clock.

If something better comes, it will be better. Sure. And we would like to have something better, because it would be better.

ch_sm•about 2 hours ago
> I tried running a smaller model locally, and it's not usable for me.

If you have the hardware, a MacBook Pro for Qwen 3.6 35B A3B and Gemma 4 26B A4B for example, they are absolutely usable, both in terms of speed and quality. Anecdotally, I can use Qwen for day-to-day coding tasks in TS and Go, without hickups.

gessha•about 2 hours ago
I’ve been experimenting with Qwen 3.8 27B and I believe I can totally use it as my main coding model provided I have the hardware for the full context. I don’t need my model to be opus level. I need it to do the tasks I want it to do without being an overprotective nanny.
embedding-shape•about 2 hours ago
I'm unable to find a local model that comes close to the effectiveness of GPT models in Codex, and I have 96GB of VRAM available and tried every local model under the sun so far. Neither of those you mention I'd say are good enough for day to day software engineering for me, but I'm also really strict about code quality and iterate on what outputs agents give me a lot before I'm happy.

With local models, this iteration cycle takes maybe 30 minutes for a single fix or feature, rather than 10 minutes with GPT+Codex, as there is so many corrections and iterations needed, although I will say that the speed I'm able to get locally makes it more fun that any of the remote models.

rapind•about 2 hours ago
> although I will say that the speed I'm able to get locally makes it more fun that any of the remote models.

This is becoming increasingly important to me. Super smart max reasoning frontier is fine if I leave it running overnight on some prepared set of clearly defined tasks, but when I want to work with the LLM, throughput really matters, and I'll go with a dumber model to get there.

At some point though, it's fast enough and any speed gains beyond that just makes me the bottleneck.

I also am seeing the smaller models gaining big strides lately, closing the gap on frontier models (still a decent sized gap though). I don't even run the small models like Qwen 3.8 27B locally. I just try them out in the cloud to see how they are progressing, and I'm definitely able to be productive.

jatora•about 2 hours ago
No, you cant. I challenge anyone who claims this to show me an actual project built only by SLM's and not using opus, sonnet, sol, or terra. Spoiler: you can't.
everyone•about 2 hours ago
You let a hiccup slip through in your comment though.
root-parent•about 3 hours ago
>> I stopped counting the number of times "estimates" said that a market would absolutely explode, and it absolutely didn't. Those are in the business of being a broken clock.

The lack of logic and risk management on this statement, is so strong, I hope humans are all quickly substituted by LLMs. Lets just do it and be done with it...

tialaramex•about 2 hours ago
> Lets just do it and be done with it...

Presumably not what you intended but this phrase immediately takes me to:

https://www.youtube.com/watch?v=dJFR7xbOIuw&t=42s

root-parent•about 2 hours ago
Great movie...yeah I think I was inspired by the scene... :-)
palata•about 2 hours ago
> I hope humans are all quickly substituted by LLMs

Why don't you go talk to your LLM instead of commenting here, then?

otabdeveloper4•about 2 hours ago
> I tried running a smaller model locally, and it's not usable for me.

Probably a skill issue on your part.

nubg•about 3 hours ago
As much as I want local and open-weights models to succeed, nothing beats a paid frontier model for now. Anybody who claims otherwise is simply not a daily user of such models. So this "investor" here should invest sime time in actually using the various LLM models and get a real taste of what it's like.
trescenzi•about 3 hours ago
Their point isn’t that local models are better or even as good more but that if you can do 50%+ of tasks with local then that’s 50% of tokens that aren’t captured as compute done in data centers.
popularonion•about 2 hours ago
> As you can see, on average, SLMs are as good if not better than LLMs in 81.2% of the cases, with the LLMs having a significant advantage only in areas like engineering, life sciences, transportation and computer sciences.

So what I’m reading here is “LLMs have a significant advantage” in the most critical areas that have practically infinite demand for more intelligence.

eigenspace•about 2 hours ago
The article is kinda dumb, and yes this is clearly the area where frontier models having and advantage matters the most, but I'd point out that these smaller open-weight models are performing better than the big Frontier models of just 4-6 months ago.

This means that the Frontier labs are under immense pressure to maintain that lead, and could end up in serious trouble if they stumble at all.

The other thing id point out is that a lot of us who are token-sensitive do things like build plans using expensive, smart models, and then execute those plans using cheaper dumber models.

Then there's the fact that we are still in the age of heavily subsidized Frontier subscriptions + tokenmaxxing initiatives from megacorps. Neither of which are sustainable, and will drive more usage to smaller open models once they end.

not_the_fda•about 2 hours ago
While that's true, the open / local models are getting good enough. Given time and the technology trend people may prefer a private local model for most use cases. Nobody is arguing that a Ferrari isn't a faster car, but the Honda is the more practical choice.
root-parent•about 3 hours ago
You completely missed the thesis here, and that is supported by the numbers being presented. It is that a large share of ordinary inference can be routed away from the hyperscalers.
hdgvhicv•about 2 hours ago
How does a current local model compare to the best frontier model 12 months ago. Or 24 months ago?
kzrdude•about 2 hours ago
It beats a frontier model from 12 months according to this bench: https://news.ycombinator.com/item?id=49334544

It is not the whole story, and knowledge is very lacking, but it has gotten a lot of attention. That model together with DeepSeek V4 Flash are the highlights of this summer on the open/local models side.

mtklein•about 2 hours ago
I have found qwen 3.8's coding quality using opencode to be similar to claude or gpt from 6-9 months ago, except much slower.