ES version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
55% Positive
Analyzed from 517 words in the discussion.
Trending Topics
#code#write#testing#tests#agents#agent#models#mutation#more#thing

Discussion (7 Comments)Read Original on HackerNews
So it's understandable that the agents wont be able to go beyond proving trival things, given how much more difficult it is to write such code.
A (more) interesting experiment (to me) would be to write a high level spec manually for a non-trivial system (liveness etc.) and see if the agent can produce an implementation using guided refinements that satisfies this specification.
Throwing tools haphazardly at the LLM and hoping they increase the correctness of its output is expectedly pretty ineffective. Good to see this borne out in the article.
The way you test code cannot (and should not) be decoupled by the way in which you architect the code itself.
80%+ of effective testing is not in the testing framework but in the code architecture.
The author doesn't mention how the code is being architected and managed.
For what it is worth, I found that forcing agents on DI/hexagonal architecture and forcing a trivial coverage check is quite useful and produce overall good enough code with relative little effort
The article claims that the agents didn’t actually use TDD or mutation testing. While it’s possible for an agent to ignore the TDD procedure even when instructed to follow it, it can’t ignore a build failure.
The whole premise of the question is hilarious. They change text in an existing one-line comment and the best models in the world think, gee, maybe I'll lint everything AND run 4000 units. Maybe 5000 integration tests too, just to ensure we collide with any other work in progress. So you write the obligatory but often-ignored obvious things into agent memory or project steering markdown or periodic nudges: You must have a hypothesis when you run expensive tests, you must spot check changes first, then start with the most relevant tests only, then move outwards only as necessary to broader labels and only then suites and only then ALL suites in a widening gyre.
But like a falcon ignoring the falconer, the models want to run the everything for anything. So you sigh, you get the model to write a deterministic hook to catch the wrong invocation of the test suite, and you spend weeks refining the rules every time you hit a edge-case, and so it goes. At least you don't have to write the involved regexes by hand, and maybe one day it will be finished..