Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
67% Positive
Analyzed from 248 words in the discussion.
Trending Topics
#market#harness#models#cost#efficiency#different#results#inference#providers#testing
Discussion Sentiment
Analyzed from 248 words in the discussion.
Trending Topics
Discussion (8 Comments)Read Original on HackerNews
Back in 2025 it was common to test models in a different harnesses.
I remember watching a guy on youtube, who was testing every new model in opencode, cline, codex, claude, etc.
Why did it come out of fashion ?
EDIT: ah, yeah. the point was that a harness would often affect results (task completion rate, I think) for more than 10%
Specifically they use this harness: https://github.com/ArtificialAnalysis/Stirrup