DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
20% Positive
Analyzed from 545 words in the discussion.
Trending Topics
#model#neuralese#cot#monitoring#english#chain#thought#monitor#don#doing

Discussion (16 Comments)Read Original on HackerNews
Or we believe the model will encode secret messages like "don't reveal this information" into the neuralese. But as the author mentions, they could have been doing that all along
> Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography
The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.
That said... CoT monitoring is a fragile "intent discovery" mechanism; neuralese puts this problem front-and-center but if agents begin to learn to hide their intent from their CoT journals, we are basically in the same spot.
If you want to monitor your model, you need to start with an inference provider that gives you the entire output and possibly even run it yourself to get access to the internal states. And if you think the KV cache and (when present) the recurrent state don’t encode a lot of “thought”, you are fooling yourself.
FWIW, I think most model architectures at least have the property that latent state can’t propagate from higher layers to lower layers by any route other than the output tokens. But even a two-iteration structure could be designed so that the last layer produces a vector that enters the first layer, once per token, and I bet it it would be very easy to train such a model to “think” in silence in the sense that the output tokens while thinking would all be one particular null token.
I think there were experiments where seemingly relevant parts of the CoT were ablated and it did not change the result.
For all we know it might be somewhat human parseable neuralese.
Or is the neuralese some sort of irreversibly encrypted data set that only an LLM can "understand" and that can never be translated back to English?
https://www.astralcodexten.com/p/elk-and-the-problem-of-trut...
https://www.anthropic.com/research/natural-language-autoenco...