Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 108 words in the discussion.

Trending Topics

#xhigh#given#opus#tasks#world#missing#frustrating#partial#presentation#apparent

Discussion (6 Comments)Read Original on HackerNews

xnorswap•9 minutes ago
A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model.

Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.

giwook•about 1 hour ago
Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?
hartator•35 minutes ago
It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.
grim_io•about 1 hour ago
I'd expect google to do well here, since they were historically strong at multimodal and physics.
jespinel•about 1 hour ago
Nice! It is missing Codex in the agent harnesses comparison IMO.
gizmodo59•about 1 hour ago
Yet another "benchmark to promote their own harness"