Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

95% Positive

Analyzed from 1082 words in the discussion.

Trending Topics

#model#models#bits#training#frontier#intelligence#every#trained#cheaper#true

Discussion (36 Comments)Read Original on HackerNews

pvillano12 minutes ago
Imagine yourself the CEO of a big AI company. It takes about a month to develop and train a model, so you release a new model every month. A startup says they can 10x your efficiency. What does that get you? You can't release a new model every three days. You can't 10x R&D either. You definitely can't tell investors that you are growing at the same rate, but selling off assets and cancelling purchasing contracts. So you just never improve efficiency enough to use less energy than the previous model version.

I don't believe this is actually happening.

awestroke6 minutes ago
> What does that get you?

Cheaper model training runs? Ability to scale training to larger model sizes without extending training time?

simonwabout 2 hours ago
> We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200.

If this holds up that's a really big deal.

wayfwdmachineabout 2 hours ago
Huge if true. As it were.
pvillanoabout 1 hour ago
A lot of people are betting their money on infinite growth forever of AI performance, compute usage, user base, subscription price.

I think cost will decrease forever.

blake__dev11 minutes ago
I agree, but I also think as AI gets better we're going to see Jevons paradox in full swing, which might delay lower costs. We've seen this with Astra according to Tibo: https://x.com/thsottiaux/status/2097559315150426222
bilaterabout 1 hour ago
Cost will keep dropping, but the frontier will keep getting pushed. The whole "give me today's model 10x cheaper and I'm good" line is a fallacy. It isn't true now and it never will be for the top 1% of tasks, which will create the most economic gains.
adam_arthurabout 1 hour ago
There are an enormous number of tasks that can get by on good enough.

If you need image recognition, and a 30B model saturates the use case with 100% accuracy, you absolutely wouldn't continue to use the next frontier model as they come out.

And I'd argue most economically meaningful tasks will be saturated by cheaper models than those requiring frontier.

Think about what today's models can do with pretty close to 100% accuracy, and then consider that they will be orders of magnitudes cheaper over the years.

5.6 Sol can already obviate tons of labor, and why would you pay 2x or more for no meaningful gain?

The relative gap between frontier and non frontier also continues to shrink, so it's not like you take a meaningful performance loss by rewinding to models from 3-6 months ago. And soon that gap will expand to 12-24 months.

I get the impression the majority of people on here only think about coding, which net net will be a tiny volume of overall AI use in the end.

bilater41 minutes ago
That can all be true but the frontier models will still have a huge market. You're thinking of all tasks as a fixed pie. The top 1% of intelligence opens up a whole new pie, stuff nobody does today because it's too expensive: daily cancer scans instead of one every few years, asteroid mining missions that need ten thousand PhD-hours of planning, custom drugs designed for your specific tumor, a personal lawyer and doctor for every person on earth, auditing every line of code in every bank and hospital continuously and so on.
saulpwabout 1 hour ago
I think there's an intelligence limit, or at least asymptote. It may be above human intelligence, but I don't think it's miles above it (at least not the kind of intelligence humans can create, recognize, or use). For example in Go, most estimates place God or "perfect play" three ranks above top professionals[0]. In the latest human-AI Go match, the human got a 2 stone handicap. So it's not like we have a lot more frontier to push there.

[0]https://senseis.xmp.net/?HandOfGod

pixl9717 minutes ago
Intelligence is spiky. In some things humans may play near the limit (Go possibly), but if you look at parts of mathematics like addition, humans can add in their head just fine but its a few trillion times faster to use a computer on addition problems of any size.

And that's not even really touching societal/network intelligence. A single human isn't that smart and can't accomplish that much. Hence we form families, and companies, and societies, and governments. What does a society of AIs look like?

bilater21 minutes ago
this assumes our whole universe and what we can do in it is a finite go board. maybe it is. but we are no where close to exploring even a fraction of it. lots of things need to happen before any limit is reached. the cavemen would also probably think we saturated tools once they saw bows and arrows.
anskabout 1 hour ago
I don't know enough about the specific models they're comparing against to say this definitively, but it looks to me like they're comparing their pre-trained models with others' post-trained models.

The metric upon which their 10x claim is based (bits-per-byte) is exactly the metric which is optimized during pre-training. Post-trained models are fine-tuned to optimize other metrics, which is known to be detrimental to performance on bits-per-byte evaluations. So bits-per-byte evaluations will always make a pre-trained model look favorable in comparison to a comparable model which has also undergone post-training.

Can someone confirm whether the models they are comparing against (DeepSeek V4, Kimi K2, and Nemotron 3 Ultra) have been post-trained?

brrrrrmabout 1 hour ago
they say they're looking at base models, so I think it's fairly compared as written.
vkaku14 minutes ago
This is great. All algorithmic efficiencies are amazing!

One thing I'd remind all scientists and the wonderful people here is this wonderful meme/line from Jurassic Park: "Your scientists were so preoccupied with whether they could they didn't stop to think if they should."

What is the actual amount of data that needs to be pre-trained and what is not? Nobody has come up with great answers to this question, and I'm already seeing amazing 0.5b-2b parameter models working very well with n-Gram corpuses of data. So, how many parameters do you really need for a given workload?

brrrrrmabout 1 hour ago
this is basically the only thing pre-training teams work on in labs. compute efficiency is the metric, the assumption that scaling = intelligence is considered a given.
ismael_rrabout 2 hours ago
Super awesome. Wish they would release the paper about what they did to achieve this. I remember nous released the token superposition paper which improved pretraining FLOPs some, but not 50x: https://nousresearch.com/token-superposition. Wondering if they also found some cool tokenization strategiesa
pixl9714 minutes ago
"Write a paper" < "Sell to a big AI lab for $$$"

Going to be interesting to see what happens to discoveries like this in the future.

monneyboiabout 2 hours ago
Imagine the sheer amount of power you could save by releasing the paper.
speedgooseabout 2 hours ago
But thanks to the Jevon Paradox, the global power consumption would probably increase.

https://en.wikipedia.org/wiki/Jevons_paradox

aaroninsfabout 1 hour ago
Solar is now not only the cheapest energy it's cheaper to intitially deploy than non-renewables.

Doesn't mean we should waste energy; it does mean that we have crossed a threshold beyond which energy concerns change shape.

pixl9713 minutes ago
"Honey, why is there a solar panel-maximizer converting our car?"
vatsachakabout 2 hours ago
Cool story. If it's true the company will be bought by open AI/Anthropic and Chinese labs will discover the trick and open source it by next quarter.
gdiamos15 minutes ago
Training improvements are very easy to copy.
vkaku17 minutes ago
Next Quarter? :) That's too long
FailMoreabout 1 hour ago
If you're like me, a SWE who is curious about ML/LLM training but unfamiliar with the terms, I got an agent to explain to me how to read the charts.

Basically, you can think of a LLM as a function which generates a probability distribution of words. If the next word in a series is "they", and one model predicts that word 40% of the time, and another model predicts that word 1% of the time, the latter model is worse as it is more surprised by the true distribution.

You can convert these probabilities into "bits":

surprise in bits = −log₂(probability of the actual token)

  Probability of actual token,Surprise
  1,0 bits
  1/2,1 bit
  1/8,3 bits
  1/1024,10 bits
This is then normalised by text length:

Bits per byte = total next-token surprise in bits / number of bytes in the evaluated text

So the lower you go on the charts, the less surprises in the LLMs distribution (a better model).

For more info: https://smalldocs.org/s/DfvdGuFsiR3LlzXw1H5J0K#k=AJ8V1AQECYj...

asadmabout 1 hour ago
not a good enough submarine