Advertisement
Advertisement
β‘ Community Insights
Discussion Sentiment
68% Positive
Analyzed from 1048 words in the discussion.
Trending Topics
#harness#context#more#memory#https#something#actually#llm#don#github
Discussion Sentiment
Analyzed from 1048 words in the discussion.
Trending Topics
Discussion (25 Comments)Read Original on HackerNews
However, currently the bigger question comes to my experience during harness is actually not where we use LLM in the system, but where we do NOT use LLM in the system. And the validation of the results becomes more and more important. Any thoughts on this?
[1] https://github.com/1stproof/batch-2/tree/main/batch-2-submis...
The idea is cool, but from own experience in harness engineering, lots of cool sounding ideas can have a negative impact on performance due to emergent and confounding effects.
So I'm a bit skeptical!
I'm much more interested in the memory model and why. As far as I can tell, it's "a vector Db" and not much more is said. Nothing about working memory or procedural memory (there are lots of ways to classify it, https://www.youtube.com/watch?v=BacJ6sEhqMo), but I was disappointed with how "advanced" it seems.
I much prefer giving the LLM a REPL loop, and injecting all the tools as functions inside the REPL loop.
That means that the LLM isn't constrained to writing a DAG, it can write code that loops, exits early, etc.
(We added the same to louie.ai, not complicated)
1. Hierarchical skills, workflow, skill learning 2. Meta Harness, self-learning harnesses 3. Trace/trajectory representation 4. Common agentic benchmarks
But first more basic things like 5. Blog posts form anthropic 6. How Claude Code/PI/ Hermes!! agent works 7. Agent sessions/ Forking/ Hooks
The example listed in the article -- fanning out a few simple get-population, get-timezone, and make-summary calls -- is, in fact, useless overengineering. This is a basic promise chain with extra steps (priced with tokens).
But as with all software pattern learning, we learn the concepts with simple toy examples that generalize into something bigger. It's the generalization that matters here.
This is talking about a few methods and tricks for spawning effective subagents (collectively, that's the "harness"). Those tips and tricks are nice, but to not be considered useless, we need to make sure we understand why spawning subagents is useful in the first place. Yes parallelism is nice for some tasks, but that's not really what this is about.
The real reason is protecting your context. Yeah, we have 1M context windows that can fit all of LotR in it, but these machines work better when they're narrowly focused. Large context windows run into attention issues and forgetfulness ("Yes, you're right, it was stated I should/n't do X but I ignored it, my bad."). So subagents come into play when you don't want all the tokens associated with a subtask to pollute your main/primary context window and degrade task attention. Split that off to a subagent, let that context navigate the details, and just make sure your main one gets just the input/output blackbox results.
The trick is getting a sense for when the complexity of the task warrants that kind of context protection, vs when a single agent is good-enough. Your toy example will never have enough complexity to warrant the setup, but you might one day find a generalization that may.
I don't think they are totally usesless.
And its clear that progression is happening on a communith level on all of these and they get integrated later on in commercial offerings like from Anthropic and co.
But also doing a opensource harness and not just giing in to the big companies allows us to have all of this open and transparent and with open models locally.
But there is another aspect which I do enjoy which is closer to the feeling of dialing in key bindings in vim or getting a really good rhythm going with your vscode extensions or zhs plugins. It is that level of "I want my system to do exactly this thing in exactly this way" customization that a lot of technical people crave.
And you can do it with harness and context engineering in many cases. In other cases it introduces friction because it will be like "cool, I will only output 15 words max unless told otherwise" and then in the next turn completely disregards it with an "oops, you did tell me to do that didn't you."
And that frustration compounds when older model versions may have done a better job of that but new models are like "thank you for your suggestion, your opinion, while appreciated, is irrelevant. Now let me get back to overspending on your token budget. "
See what some guys like Linus Torvalds, or Eric S. Raymond are saying about. It's not so much about "vibes" but using the tool (yes the AI tool) in a certain way that can propel yourself towards your goal at unprecedented speeds.
Trying to make gold from pyrite.