DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
81% Positive
Analyzed from 979 words in the discussion.
Trending Topics
#models#model#bash#harness#better#capable#tools#harnesses#success#different

Discussion (15 Comments)Read Original on HackerNews
Basically if a Car A is performing better (be it speed, milage or in general sense) than Car B, then it is not necessarily because its engine. It could be because of better tires, better gearbox, lighter body, better usability of features, etc.
You can implement an AI feature (like AI for BI) in different ways even with the same model - via ReAct-loop, or plan-and-execute, or hybrid. You can make it stateless, stateful, RAG-based, etc. depending upon whether you want to prioritize result accuracy or depth of analysis. You can use LLM to generate either intent (requires lesser reasoning) or the queries itself (requires much more capable model).
Your harness can adapt to the underlying model's native capabilities, or can make up for its absence, e.g. query generation in above example requires your model to have MOE capabilities but intent generation wouldn't.
And there might be a point to these arguments, vaguely. However:
There never seems to be - any - kind of counter example or reasoning behind the rationale. You have an in depth and empirical study, done by researchers who, frankly, now their shit (most of the time)
And on the other hand a random internet comment saying "nope" because...the models aren't the latest.
If the latest models really would make a difference, you should at least provide some kind of evidence towards that. As it stands though, every time these comments come up this is missing.
There seems to just be a vaguely defined understanding that "everything changes all the time, and nothing you ever research is transferable to state-of-the-art models"
Which brings me to my second point about these kinds of arguments:
LLM models often - aren't - fundamentally different. Yes, they are vastly more capable. And yes, there are emergent properties. But at their core, they function very much similarly. And for quite a while now, there have not been any of these drastic changes we saw when LLMs first become "good enough" for agentic coding.
I am tired of dismissing empirical evidence and studies every. single. time for reasons without evidence and seemingly a vague sense of "no, but my model is different"
> Planning improves success at additional cost for weaker models but mainly reduces cost, with small decreases in success rate, for stronger models.
> Predefined tools raise success rates for models with weak bash control, whereas bash-only yields higher success at lower cost for bash-capable models, most clearly on shell-centric task types.
> context management extends execution trajectories without substantially altering agent behavior and is most beneficial under tight context budgets
> planning sustains the trajectories of models that abandon tasks too early and trims repeated verification in models that verify too long
> structured tools support models with limited shell proficiency, while bash-only enables capable models to combine multiple code modifications in a single tool call
Seems fairly intuitive to me, based on feeling. But also fairly kind of obvious; bash-only tooling has higher success for bash-capable models, compared to using predefined tools for models that aren't good at bash? Yeah... They all seem a bit "duh" to me. The final piece of the conclusion is agreeable regardless of how they arrived at it though:
> Harness design is thus a conditional systems problem in which each component should be selected for the target model, task type, and resource budget rather than adopted as a default.
I think lots of people treat the harness/model/prompts combo as interchangeable, but in my experience the quality and efficiently depends heavily on the combo of the harness/model, and using the harness + model made by the same lab, has vastly better experience compared to more "general purpose" (for the lack of a better term) harnesses. Most likely because they use their own traces when training future model iterations.
No. The conclusion is that:
bash-capable models + bash-only tools > bash-capable models + predefined tools
In other words, MCP was just a bunch of bullshit that maybe helped a little bit until the models got good at bash, and now it's basically useless.
"Duh, of course Mars has canals."
Testing the "obvious", "duh" things is incredibly valuable science. It provides a more solid foundation on which to build because it reduces the assumption space.
not really, there was a recent benchmark with claude and codex and it showed no difference in ability with a harness like pi agent compared to their native harnesses, pi was in fact cheaper per task.
Anthropic appears to agree frontier models don't need in-session planning tools.
https://github.com/anthropics/claude-code/issues/80487