ZH version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
86% Positive
Analyzed from 284 words in the discussion.
Trending Topics
#verification#more#right#pixel#weeks#claude#better#agents#thing#prompt

Discussion (7 Comments)Read Original on HackerNews
> The verification is probably the single most important thing that people do not get right, largely.
and so, so, so wrong about giving this prompt (for verification) and expecting that it succeeds at building a "good" app:
> I want you to run the Electron app in the Mac virtual machine, screenshot it, and then look pixel by pixel. Compare it to the Swift version. Don’t stop until you’re done.
Given he has let it rip for two weeks, I am assuming they have been post-training Claude for some version of this to be more _likely_ successful than not. However, IMO, the verification that you get from a visual comparison is shallow.
To state the obvious: there is a lot more than what meets the eye. But, I think, one could prompt a Fable/Opus 5 to actually go verify that "lot more"...
The question is: should one be imperative in asking for a specific types of verification (like a rubric) vs hoping that the Google/Anthropic/Open AI/Moonshot's post-training will take care of it.
I think, as things stand today, even with the best-in-class models today, I would be leaning more imperative. And it is not because I am an expert in SwiftUI or such. It is because I want to be able to say that _I_ (i.e. the human) verified that this thing works.
Agents are writing a mountain of code to prove how good agents are, and then humans throw it away.
Next time, on Dragon Ball Z!