ZH version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
67% Positive
Analyzed from 248 words in the discussion.
Trending Topics
#market#harness#models#inference#providers#testing#common#test#different#why

Discussion (8 Comments)Read Original on HackerNews
Back in 2025 it was common to test models in a different harnesses.
I remember watching a guy on youtube, who was testing every new model in opencode, cline, codex, claude, etc.
Why did it come out of fashion ?
EDIT: ah, yeah. the point was that a harness would often affect results (task completion rate, I think) for more than 10%
Specifically they use this harness: https://github.com/ArtificialAnalysis/Stirrup