Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 71 words in the discussion.

Trending Topics

#models#benchmark#expects#happy#path#edge#cases#though#ticket#model

Discussion (4 Comments)Read Original on HackerNews

andhuman1 day ago
This benchmark expects models to not only do the happy path, but also edge cases, even though the ticket they give to the model doesn’t specify if. This to align more to real world tickets. Usually with benchmarks it’s the other way around: only implement what’s asked. So I welcome this type of benchmark because this is how I use the models.
klooney1 day ago
Is this a proprietary harness? I feel like harnesses have a huge influence on how models behave
chalmovsky1 day ago
chalmovsky1 day ago
this seems really low?