HI version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
53% Positive
Analyzed from 1481 words in the discussion.
Trending Topics
#intelligence#agi#test#human#more#game#don#humans#benchmark#arc

Discussion (49 Comments)Read Original on HackerNews
I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we used the Turing test as the proxy for AGI, but then early LLMs could clearly pass for a human in a casual conversation while clearly not matching human performance on most other tasks.
Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.
It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.
So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.
This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs
I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations
[0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...
https://arcprize.org/
Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.
Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."
Well I dont know about all of you, but I am celebrating meat based humans...
Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"
Examples:
- predict a coinflip: easy to verify, hard to learn
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
but we both know these examples go against the spirit of my point
Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.
Alignment++
The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).
Prediction:
We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.
TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".
https://www.lisep.org/tru
(I have not gone down the rabbit hole to understand how they achieve that 24% number)
IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!
It's already happening :)