ZH version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
36% Positive
Analyzed from 1128 words in the discussion.
Trending Topics
#model#neuralese#english#monitoring#models#cot#don#doing#thought#more

Discussion (27 Comments)Read Original on HackerNews
Or we believe the model will encode secret messages like "don't reveal this information" into the neuralese. But as the author mentions, they could have been doing that all along
> Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography
The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.
It's really hard to definitely prove it's not doing it right? Hopefully the model does not do anything like this during training because its too much work
So you can turn it to English, but only to a LOT of English, and doing so would slow the model down a great deal, and it would be a lot more like a detailed thought than a sentence.
We’d have models monitoring models as our only way to know what they’re planning.
A great movie on this is “Collosus: the Forbin Project”. Shot decades ago. The computers discover the other computers and start communicating — and bootstrap their own language — much like we saw happen with OpenAI agents.
https://www.reddit.com/r/scifi/comments/1nl4vex/colossus_the...
If you want to know what a simple version of Neuralese communication looks like, look no further than Facebook’s Marketplace agents experiment a couple years ago.
And all that was actually constrained by English and the FFN
They published open weights versions of these interpreters for a number of open models sometime in the last year. Very cool idea.
By the way, they concluded CoT often lied, based on the neuralese interpretation.
EDIT: a comment below linked to https://www.anthropic.com/research/natural-language-autoenco..., which is what I was referring to.
That said... CoT monitoring is a fragile "intent discovery" mechanism; neuralese puts this problem front-and-center but if agents begin to learn to hide their intent from their CoT journals, we are basically in the same spot.
So, you mean, like another human person?
https://nonlineartransform.substack.com/p/relax-about-neural...
There's also an argument here for why its _better_ for monitoring (because we have the whole state space).
https://www.astralcodexten.com/p/elk-and-the-problem-of-trut...
It discusses "Eliciting Latent Knowledge" which is "a technical report / contest / paradigm run by the Alignment Research Center". The research investigated whether it would be theoretically possible to build a "trustworthy" AI to interpret the thoughts of another AI.
Take anything they write with a big grain of salt. EA writings these are mere apologies. The conclusion is preordained. Authors start with the goal of slowing AI and work backwards from there, trying to see which arguments resonate with the pubic. You can't unsee it.
Not everyone in the AI space approves of these people or their doomerish.
I think there were experiments where seemingly relevant parts of the CoT were ablated and it did not change the result.
For all we know it might be somewhat human parseable neuralese.
https://www.anthropic.com/research/natural-language-autoenco...
If you want to monitor your model, you need to start with an inference provider that gives you the entire output and possibly even run it yourself to get access to the internal states. And if you think the KV cache and (when present) the recurrent state don’t encode a lot of “thought”, you are fooling yourself.
FWIW, I think most model architectures at least have the property that latent state can’t propagate from higher layers to lower layers by any route other than the output tokens. But even a two-iteration structure could be designed so that the last layer produces a vector that enters the first layer, once per token, and I bet it it would be very easy to train such a model to “think” in silence in the sense that the output tokens while thinking would all be one particular null token.
Or is the neuralese some sort of irreversibly encrypted data set that only an LLM can "understand" and that can never be translated back to English?
https://www.astralcodexten.com/p/elk-and-the-problem-of-trut...