HI version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
100% Positive
Analyzed from 263 words in the discussion.
Trending Topics
#harness#self#code#agent#https#interesting#run#tried#using#engineering

Discussion (8 Comments)Read Original on HackerNews
Curious if anyone's tried using RL for harness engineering? I think we're still pretty far away from the optimal harness, especially when it comes to long-context memory management.
I guess it depends on the model you're trying to use, but seems most of them prefer smaller codebases, they work a lot better with less code, which kind of makes sense. With that in mind, I'd probably aim for something way smaller to bootstrap a self-improving agent. Then I'd use this "Prime Agent" as an example to my self-improving agent for what it should not evolve to.
I am curious - how does it fare for other benchmarks, or everyday programming?
Is it that it wasn't accepted yet, or are there issues with how it was run?