FR version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
100% Positive
Analyzed from 237 words in the discussion.
Trending Topics
#metr#test#arc#agi#interesting#claude#measures#different#trackingai#humanity

Discussion (6 Comments)Read Original on HackerNews
- METR's time horizon
- TrackingAI's offline cognitive test
- Humanity's Last Exam, and
- ARC-AGI-2
that have lasted longer than 2 years (though ARC-AGI-2 is now saturated).
When plotted in the linear domain, they all have an exponential (hockey-shaped) curve, but the interesting thing is that the bend happens right at Q4 of 2025 (right when Gemini 3, Opus 4.5, GPT 5.2 all come out).
I do use deep research (across Gemini, ChatGPT, and Claude) to gather background and ideas, and Claude for editing. I started machine learning and computer vision research back in 1995, studied NLP, and dove into LLMs with GPT-2, so I wouldn't be surprised if my writing has been deeply influenced by AI.
Longest? Yes, they actually use actual wall clock time. I don't think this metric makes any sense. I don't know about the others, I stopped reading at this point.
What's interesting is that while they are rather different in nature (yes it is odd that METR measures clock time as opposed to iterations etc.) but the behavior of the resulting improvement curves are extremely close.
That's the punchline: 4 independent measures point to the same conclusion. That increases the chance that the conclusion is correct.