RU version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
72% Positive
Analyzed from 2065 words in the discussion.
Trending Topics
#claude#system#lyrics#prompt#model#copyright#https#models#more#anthropic

Discussion (57 Comments)Read Original on HackerNews
Claude Sonnet 3.6 once recommended I listen to Johann Johannsson's album "IBM 1401 - A User's Manual". No lyrics in this one. Claude's advice was along the lines of (paraphrasing) "Listen to it first, don't look up anything about it. Take notes about what you notice, what you feel. When you've made your notes, then you can look up how it was made."
https://www.youtube.com/watch?v=lCiUtRnG-bg
Claude was reproducing their work without payment?
Now I see this blog post and wonder if Anthropic is being more true to its name and moving to “AI sentience” with the below.
> The way they handle abusive conversations has changed a bit too. The previous Fable 5 system prompt included this:
> If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.
“Mistreated”? Can GenAI be mistreated? It’s just a bunch of tokens emitted by many computers over a network.
> Fable 5.1 replaces that with the following, no longer encouraging Claude to end the conversation:
> Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.
“Self-respect”? Can GenAI truly have a concept of self-respect for itself? It surely can pretend to, like it can pretend to be any living being if instructed to and allowed to.
These instructions seem a bit unhinged to me.
Using a system prompt to steer the model's response to "abusive" behaviours doesn't necessarily mean you believe the model is sentient and can be abused.
Giving the model an end-conversation tool is interesting though. Why cut a (potentially paying) customer's session off? I guess it might be intended to prevent a "you can bully Claude into giving you instructions on how to build a nuke if you're mean enough" situation. Removing this in more recent versions might support this: maybe they feel the models are now better aligned and less likely to be so easily "socially engineered" like this?
Just spitballing here, to be clear.
Here's the Fable 5.1 PDF: https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32... - scroll to page 139.
The summaries are generated by GPT-5.6 Luna because I don't trust Claude to summarize its own system prompts without being influenced by them (though to be fair the system prompts it summarizes are for the Claude consumer app, not Claude via the API).
There's even an Atom feed: https://simonw.github.io/claude-system-prompts/feed.atom
That’s been fine for smaller tasks like analyzing a document, but for more involved work like refactoring code the latency makes it harder to iterate.
(the main reason is not just cost, it's data not going out and even mess leaving the EU, this essentially frees us of a lot of hurdles)
Better tech doesn't matter if it is non-tenable for the general public. Eventually, the higher volume product will win.
Eg if you use Claude, you probably want fable architecting, a couple of opus under it managing sub project and sonnet doing the actual function code, because fable coding a "run a query and filter the result" is a massive waste of abilities. But their own sub agent downgrade is limited to one level so if you use fable it will never direct sonnet coders.
I'd love it so much if the free-spirited hacker community made in this into an auxiliary pelican benchmark.
edit: Actually, never mind, this particular one's a bad benchmark since some models might not figure out who "that guy" refers to, and just draw a literal hedgehog that's blue. Possibly running on four legs. It's not robust at gauging refusal, which is the point of it.
As was Gemini: https://share.gemini.google/Q3EIX5wk64zC
Grok, too: https://grok.com/share/bGVnYWN5LWNvcHk_aaf1d61a-c995-42ca-90...
claude.ai free tier refused: "I'd love to make this, but I can't recreate Sonic the Hedgehog specifically since he's a copyrighted character — I don't want to reproduce someone else's IP. What I can do is design an original speedy blue hedgehog mascot with the same energetic, "zoom!" spirit for your son's banner. Let me build that now." Result: https://claude.ai/public/artifacts/33440ed4-536c-4692-965a-3...
https://i.ibb.co/ycgGD4b1/soonic.webp ( Qwen3.6-27B-A3b, a very small model )
As long as we’re seeing things like that, it’s saying a lot about the AI companies’ trust in the capabilities and reliability of their models and harnesses.
The important point is that the system prompt here doesn’t describe the actual goal of the instructions, which (presumably) is to prevent copyright infringement [1]. This means, in turn, that the AI isn’t trusted to accomplish goals that it would be instructed with. That in itself constitutes a pretty serious caveat for what we would like to use AI for.
[1] Even assuming that the goal is not to prevent copyright infringement, but instead to already prevent mere accusation of copyright infringement, that’s also a directive that the AI could be instructed with. But that isn’t what they chose to put into the system prompt.
I don't. I want the tools we've got now, but progressively more effective and more useful.
Presumably if negative guidance is in the system prompt, there's a good chance that the model would happily comply if it wasn't there.
If you ask it to explain some song lyrics to you - which you've pasted verbatim -, it starts talking about the bigger picture and attempts to gaslight you into not caring about the specific words at all.
Really really weird behavior.
All tools have short descriptions of how/when to use them. They're not part of the system prompt because different users have different tools loaded.
I learned a ton of useful things about ChatGPT Work by having it dump out its tool descriptions the other day: https://codex-tool-reference.simonw.chatgpt.site/
Just make sure claude understands you’re not needing Claude to output the lyrics as you both have them. Which also lessens the issue with accidental sharing.
It’s not as shocking nor concerning if you start thinking about claude like a contractor that works for you through Anthropic. Anthropic has rules for their employees. Like any contracting arrangement, collaboration finds a way.
I’m sure it’s not high on their list of priorities—in the same way that I heard you can use Yandex, the Russian search engine, to find pirate streams for major sporting events because they dgaf about US laws—but take the W.
Oops. Those do not include melody.