Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 237 words in the discussion.

Trending Topics

#metr#test#arc#agi#interesting#claude#measures#different#trackingai#humanity

Discussion (6 Comments)Read Original on HackerNews

rbuccigrossiabout 17 hours ago
Author here. In short there are 4 different metrics:

- METR's time horizon

- TrackingAI's offline cognitive test

- Humanity's Last Exam, and

- ARC-AGI-2

that have lasted longer than 2 years (though ARC-AGI-2 is now saturated).

When plotted in the linear domain, they all have an exponential (hockey-shaped) curve, but the interesting thing is that the bend happens right at Q4 of 2025 (right when Gemini 3, Opus 4.5, GPT 5.2 all come out).

Khaineabout 18 hours ago
It reads like it was written, or heavily edited by AI
rbuccigrossiabout 17 hours ago
The text is my transcript of the video formatted by Claude and with sources added at the end.

I do use deep research (across Gemini, ChatGPT, and Claude) to gather background and ideas, and Claude for editing. I started machine learning and computer vision research back in 1995, studied NLP, and dove into LLMs with GPT-2, so I wouldn't be surprised if my writing has been deeply influenced by AI.

Khaineabout 14 hours ago
Ah, makes sense. Its well written and an interesting piece of work, I just noticed hints of 'claudisms' in places.
ffaccount2about 21 hours ago
>METR measures the longest software engineering task a frontier model can complete successfully half the time

Longest? Yes, they actually use actual wall clock time. I don't think this metric makes any sense. I don't know about the others, I stopped reading at this point.

rbuccigrossiabout 17 hours ago
The others are an IQ test (TrackingAI), a test of Ph.D. questions across multiple domains (Humanity's Last Exam), and graphical pattern matching (ARC-AGI-2).

What's interesting is that while they are rather different in nature (yes it is odd that METR measures clock time as opposed to iterations etc.) but the behavior of the resulting improvement curves are extremely close.

That's the punchline: 4 independent measures point to the same conclusion. That increases the chance that the conclusion is correct.