DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
63% Positive
Analyzed from 943 words in the discussion.
Trending Topics
#agi#harness#gpt#don#nvidia#models#avo#arc#https#opus

Discussion (31 Comments)Read Original on HackerNews
https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-...
There you will find the extremely important qualifier it's the public set, not the private set (with the risk of overfitting, ie the results not repeating when submitted to be run in competition), and the detail that this is essentially a harness added to Opus 5, not Nvidia's own models.
Obviously still impressive, you would think.
GPT-4 was decidedly not capable of beating Pokemon 18 months ago. I doubt it would be able to complete a single level. I don't think people realize how large the advances in model capabilities have been. GPT-4 in a modern harness is absolutely horrendous compared to modern models.
Have you ever played pokemon?
True as that may be, it may be better to optimize models for some amount of memory versus forcing some token count based on a reasoning level, right?
Most of this doesn't discredit your overall point, though.
Community maintained spreadsheet of the runs: https://docs.google.com/spreadsheets/d/e/2PACX-1vQDvsy5Dt_-P...
What NVIDIA has here is a generic "evolution" harness, which can be used for any problem.
I think it would be fair game to allow OpenClaw, Hermes, Codex, Grok Bot, this NVIDIA thing, to compete, as long as they don't have ARC-AGI specific skills, toolset.
Using Claude Opus 5, but it can use others:
AVO is also designed to operate across frontier models. While our full public-set result used Claude Opus 5, we additionally paired AVO with GPT-5.6 Sol on a challenging subset of games. In these limited experiments, Sol reached matched levels faster in wall-clock time in several cases, while Opus used fewer environment actions in matched-level comparisons. These preliminary results suggest complementary operating profiles across models, and we leave a broader systematic comparison to future work
Curious to read more about it though, seems the paper for it is here: https://arxiv.org/pdf/2603.24517, I'm not sure I understand if it's better than just Codex with a /goal, as they talk about "can discover performance-critical micro-architectural optimizations" but leave Codex alone for a day or two and you'll get the same results without doing "additional autonomous adaptation" at all.
Anyway yes I think we've had AGI for a while now, even if the GI doesn't quite match up with what we expect from a human.
100% is some "RHAE" metric: its performance of median human first time seeing those problem.
Yey, AGI is finally solved.
(0) https://arcprize.org/arc-agi/3
https://arxiv.org/html/2603.24517v1
The next year is going to be wild folks