Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 198 words in the discussion.

Trending Topics

#agi#site#benchmark#score#published#every#model#less#benchmarks#human

Discussion (1 Comments)Read Original on HackerNews

baraklaniadoabout 3 hours ago
I run AGI Ranker, an AI benchmark aggregator that calculates an AGI score for every benchmarked frontier model, and have just released v2 earlier today. Auditing my own scale has led to a lower AGI score, 6-15pts less, yet the models kept their shape. Only one of the ten benchmarks has actually measured a human ceiling, GPQA-Diamond 81%, so the other 9 are anchored to benchmark-max, not to human-parity. Coverage has greatly improved, from ~43% to ~80%, and that means the score leans less on shrinkage and more on measurement. Single evaluator dependency has been successfully dropped from ~45% to ~26%, with the new offsets published on the site. Agency was rebuilt on 4 benchmarks from 3 independent evaluators. Speciality tabs now include Coding and Knowledge, with Reasoning temporarily dropped because it was leaning on only one benchmark, but will make a comeback once AIME 2026 is published. The Value tab shows cost-per-capability. The Corrections log makes sure that all my mistakes are openly reported on the site, including a recent apology to Deepseek, for having published an unsourced ARC-AGI-2 number. No lab money, no paid placements, every cell traceable, open data (CC-BY). It's a solo project that I maintained while having only 450 monthly visitors, determined to offer those who do visit, the most accurate and useful information about AI model prowess. Known limitations are listed on the site - happy to answer questions about the methodology or the functionality of the site. Cheers!