Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

50% Positive

Analyzed from 119 words in the discussion.

Trending Topics

#index#sol#scores#gpt#astra#lower#title#major#gains#chart

Discussion (8 Comments)Read Original on HackerNews

NiekvdMaasabout 1 hour ago
Title: "major gains"

First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)

flyaway123about 1 hour ago
Indeed. Though to be fair it is referring to "Artificial Analysis Coding Agent Index", from 65 to 67.
dist-epochabout 1 hour ago
I think they mean cost per task, where Astra is now on the Pareto frontier.
x3haloed20 minutes ago
Fascinating. This is the only benchmark I've seen so far with lack-luster results. I don't understand enough about AA's specific methodology to get the implications.
ekojsabout 1 hour ago
Well, seems like ECI [0] and the AA index is diverging quite a bit. Benchmarking LLM is tough and I think we are seeing the limitations of current benchmarks and applicability to real tasks.

[0]: https://x.com/EpochAIResearch/status/2095602754282783108

eis40 minutes ago
In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58). In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3.

Am I missing something or is this not looking too... stellar?

Readeriumabout 1 hour ago
more like 5.7 not 6
lostmsuabout 1 hour ago
5.6.1