FR version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
66% Positive
Analyzed from 3439 words in the discussion.
Trending Topics
#opus#models#claude#model#more#code#prompt#better#user#prompting

Discussion (93 Comments)Read Original on HackerNews
A common tactic is to used a big brain model like Opus for planning and reviewing, and a cheaper model for execution.
For regular software development they have been pretty great.
Non-pedantic answer: I totally agree with you. Opus 5.5 is totally knocking it out of the park IMO.
Which was and is true to some extent.
And don't get me wrong, China is a dictatorship, and a tyranny for some.
But then again, the west is a tyranny for some.
Doesn't make it any less amusing from the outside, to see the US struggle with their identity. (It's most always just a struggle when freedom becomes less)
AI is rapidly saturating it's ability to be useful and these products need to start to mature.
It's not 'fun' to manage 50 different broken MCPs and their variety of ways in which they are broken.
It was 'fun' at the start, now it's just 'broken technology'.
Astra and Opus 5.5 are the 'starting point' for the next era of AI where we expect robust tooling.
The reason why advanced prompting is a moving target is that a lot of prompting is "use extra instructions to compensate for specific ways in which the target LLM is weak or prone to errors". And guess what? LLMs get better over time - obsoleting your advanced prompting.
"Tune a prompt to death for the specific task and specific model" gets you better performance in the moment, but "trust LLM to be smart" ages a lot more gracefully.
I'm genuinely worried about all our short term investment in mitigating the failure modes of models that may only be SOTA for a few months.
It's very possible people being 'late' adopting AI may end up with a leg up, not only because they spent more time polishing personal skills during this time, but also because they don't bring all the baggage of 'AI competence' that is becoming irrelevant at breakneck speed.
That could be seen either as early adoption that’s overfitted to current capabilities or as late adoption of LLM’s more advanced capabilities.
It was even more 'broken' at the start. We overcame some of the issues by 'prompt engineering', which is needed less in the newer, smarter models.
The first combustion engine was a miracle. It only becomes 'broken' when we evaluate in some kind of applicable context.
There's so many "x generated this in one shot, this is agi" stuff that gives you the impression that you can vibe operate modern models the same way you operated last year's models. There's so much more to it than that. It requires you to put a faith in the leap in the capability of models, one that would've surely been a waste of time in previous models.
Not sure where i'm going with this other than I think most can relate that it's exhausting keeping up with. I cant imagine what it'd be like parenting a kid that went from toddler to puberty in the span of a year and planning for them to go to college the next year. This industry is moving so fast that it's becoming fact that it's the user that's "holding it wrong" every six months.
The step function change on Opus 5.5 for visual work shocked me.. and I haven't been surprised like this in a long time with LLMs.
EDIT: When I first saw the "P(DOOM)" video and some of the other animations I was VERY skeptical that Opus 5.5 without a lot of tools could make something like that.. until I tried it for myself. It can.. 100%.
But it other cases, like the music videos, much of the magic is done by access to elevenlabs and suno apis.
Edit: just saw your edit about the pdoom video. Can you share how you prompted it? Would be helpful to know.
I'm still more worried about the malice and any malicious acts by the people at these frontier labs than the models at the frontier labs.
Yet here we are, "why my calves hurt more than any other muscle after training" being classified as a naughty question.
Opus 5.5 has been amazing, but I'm confused by how this is worded. It "matched or beat" Opus 5? There is no matching. There is only surpassing. By miles. Like Opus 5 was the biggest disappointment of the year. Opus 5.5 is even better than Fable. I do not understand why they're not acknowledging it for the leap that it is?
The data doesn't support it being better on every test (sometimes the score will be the same imperfect one, sometimes both will have gotten a perfect score).
I feel like hybrid AI-driver UIs are a bit underexplored and are probably a good way to increase visibility. Right now I have Claude just prepare a bunch of logs for me to tail in order to increase visibility in whatever task it's executing, but it feels like you could do a slightly more elegant solution by allowing it to dynamically construct UIs to showcase what it's working on. Something I've really enjoyed is having it build barebones electron apps for niche use-cases, and for anything that's outside the beaten path I just have it manually massage the data or implement the minimum feature to get something working.
Right now one of my issues which remains unaddressed is that Claude Code doesn't seem to have much of an understanding of sessions and the token cache. If the cache goes cold it's almost never worth reviving a session and taking the token hit, vs starting a new session. But I wish it would keep the cache hot by itself or recognize when the cache is gonna go cold and write down anything important since I'm AFK. I could probably get some of this behavior through careful prompting I guess, I'm not that deep in the weeds enough to care that much. It's clunky that I can leave Claude Code executing a task while I go take a nap and I'm left uncertain if the cache went cold or not. I'd really like a gated "Are you sure?" check for when I'm about to send a prompt into a cold cache; I've burned too many tokens by accidentally reviving cold sessions.
Is this a problem with the model or the harness in your opinion?
I queried it and was told that sub-agents can't run processes a blocking fashion, I'm not sure the harness changed, or the model was handling it differently, but it require some changes to skills to prompt around it.
Of course, and this is the basics anyone should do when working with LLMs & agents; but with their high-variance, doing statistically significant benchmarking is very costly. Which is why the debates here on HN often talk about the "feelings" of degradation (or improvement!), but often without proofs. I'm not sure how to solve ạt; maybe inference providers should provide free benchmarking to anyone publishing results, along with the guarantee to never train on those sessions.
Claude Code has an output style setting that I set to "Concise", with no apparent effect.
I am told this is merely something in the system prompt that the model tends not to pay attention to with large contexts.
Opus 5.5 writes whole essays at the end of the turn, with the important actionable steps somewhere at the bottom.
When prompted to give a concise summary, it usually overshoots into a super short summary and then you have to dig into the details again anyway.
In general I find Opus 5.5's writing to still have more "ticks" or "Claudisms" than the OpenAI models.
Its explanations often appear overcomplicated for simple concepts.
Sure, it's leagues above the ridiculous writing of Opus 5, but Anthropic still has a long way to go here.
IIRC it's a system reminder injected after every single turn.
It must be pretty ingrained to be so resilient against prompting. I think RL on relatively short-horizon programming tasks has given the model a tendency to write down absolutely everything, so it survives compaction. Longer-term (project-scale) tasks where this crap starts to pile up and cause problems are in the evolutionary shadow, so to speak.
I guess I could take some lengthy example explanation, and have it try various instructions and test what results in output that I find preferable.
Maybe I'll give that a try, thanks!
I'm not sure I understand this complexity. In all harnesses I've ever used, tool calls themselves are surfaced to the user as an indication of progress. When the UI/UX around this is engineered well, the user should be able to infer roughly what is going on. Different tools have different ideal presentations. You can't reduce everything to plaintext blobs.
If I absolutely needed intra-turn progress updates, I'd accumulate a separate per-turn transcript and feed it into a cheaper model at deterministic intervals.
In the past I've been very skeptical of this kind of protection. Anthropic have clearly trained their models for this though, so maybe Opus 5.5 is smart enough for this to work?
Will be interesting to see if minds more devious than mine can break it.
Well, maybe? There is a lot of valid XML ingested in the training data, so I wonder what happens when the model encounters:
I used Opus 5.5 for some simpler tests and was quite angry when I saw that each of my question was above 10USd
In which case we've royally fucked ourselves that the level of engineering we've reached is... prompts. Because there is a deadline where we have to show productivity to justify all the investment spending.
People need to build with tools in a reliable, constructive way. Not vodoo magic based off vibes. We need better structured output, better transparency on what these models can do, better controls overla, maybe new ideas on loops graphs, and ways to use the models. Like, at least people were trying new things with jev.
I say the above because I'm seeing entire worlds and games being one-shotted built on X and I just have no idea how they do it. I tried building a large prompt for Fable when it was first released and it didn't have anything close to resembling some of the stuff I'm seeing today.
Lies?
> Asked for frontend work without design direction, Claude Opus 5.5 falls back on a few default styles, and a general instruction such as "avoid a generic AI look" mostly swaps one default for another. It responds well to instructions that name specific patterns to avoid, as in the following example. Work iteratively: check which styles the first result used instead, and extend the list if needed.
I hardly ever read tips for prompting etc. because things change too quickly, the writeups are kindof big. Glad I read this one, because I often did exactly what they assume users would do. I write "don't make it look like generic ai slop" and that seemed to work nicely. Now I know why there was still a chance of seeing similar styles across apps. I reckon doing some manual work in terms of scouting dribbble/behance for nice layouts will yield better results.
The point is, it doesn't. As the prompt says, there are a few default styles, and without design guidance the model just chooses one at random:
Asked for frontend work without design direction, Claude Opus 5.5 falls back on a few default styles, and a general instruction such as "avoid a generic AI look" mostly swaps one default for another.
In other words, in response to "avoid a generic AI look", enforce the exact opposite of that user prompt and literally choose a generic AI look. Which I must admit, I kinda love. Meet low effort prompting with low effort results. "Oh, you didn't like this generic style? Try this other generic style on for size. You're gonna love it!"
(But why only "mostly swaps"? Is Anthropic letting an occasional lucky user hit novel AI design gold?)
Cars were just like that during their early years, with tillers and knobs. See the video where Top Gear finds the first car with controls we recognize https://www.youtube.com/watch?v=fkwGJzU5B-I
You can stay on your horse. It's perfectly usable. Don't fall for they hype.
Edit: Someone commented that this is insulting to autists, and I guess it kind of is - sorry. What I ment was an intellectual; one that can be an absolute retard, but have read an aweful lot.
Not sure how you missed it but that's exactly what I'm calling out as asinine.
> We're in a race right now
Again: consider the analogy about cars... which are literally used for racing (occasionally).
i dont understand how you're framing this. how is this a bad thing exactly? How is it asinine?
Just saying "continue" when it gets stuck usually makes it repeat the same error. A better way is to save its last action and result, then make it try a new approach. If it tries the exact same thing twice, it should stop and ask the user for help instead of wasting money on a loop