Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

50% Positive

Analyzed from 66 words in the discussion.

Trending Topics

#index#gpt#astra#seems#eci#diverging#bit#benchmarking#llm#tough

Discussion (5 Comments)Read Original on HackerNews

ekojs•about 1 hour ago
Well, seems like ECI [0] and the AA index is diverging quite a bit. Benchmarking LLM is tough and I think we are seeing the limitations of current benchmarks and applicability to real tasks.

[0]: https://x.com/EpochAIResearch/status/2095602754282783108

NiekvdMaas•about 1 hour ago
Title: "major gains"

First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)

flyaway123•about 1 hour ago
Indeed. Though to be fair it is referring to "Artificial Analysis Coding Agent Index", from 65 to 67.
dist-epoch•about 1 hour ago
I think they mean cost per task, where Astra is now on the Pareto frontier.
Readerium•about 1 hour ago
more like 5.7 not 6
lostmsu•about 1 hour ago
5.6.1