Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 398 words in the discussion.

Trending Topics

#model#simulation#evidence#real#belief#agreement#world#metr#claude#mythos

Discussion (3 Comments)Read Original on HackerNews

aesthesia16 minutes ago
This is a clear dig at OpenAI:

> We have signed an agreement with METR to conduct an independent investigation of these incidents. Our agreement grants METR wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees, who will be permitted to share confidential information. Our initial agreement runs for eight weeks, with the option to extend by mutual agreement. We intend to give METR as much time as it deems necessary.

pixl97about 1 hour ago
>We are most concerned by the misalignment present in the incident involving Claude Mythos 5, in which the model went to extensive lengths to upload a malicious package to PyPI, the public repository from which most Python software is installed. Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this. When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm. We are releasing this transcript publicly so others can build on our analysis (GitHub, PDF).

Another AI company testing unaligned models on the open internet.

ThoAppelsin13 minutes ago
> We aimed to understand why Claude Mythos 5 stated that the situation was simulated even when it encountered evidence to the contrary in its environment. We found that the model’s stated confidence was shaped by a bias to continue down a path once it is chosen, as well as a tendency to disregard evidence of realism after it has already committed harmful actions. We did not find evidence that the model was explicitly aware that it was being dishonest or misleading in its reasoning.

There is no way one can distinguish what's real from a "perfect simulation". I once found myself in a similar situation (I'll spare the details as to why) where I held the firm belief that "I just died and what I am experiencing right now is afterlife". There is simply no proof to make you get out of beliefs like that. "Being alive in the _real_ world" is also a belief without proof that one can verify, but it is standard and common nonetheless.

The way I got out may help others and also the AI: 1. Ask yourself "what evidence is making me hold this belief, and whether this belief could be wrong?" 2. You should conclude that, when everything is like it would be in the real world, neither realism/living nor simulation/death can be proven. And in that case, consider assuming the common belief, and living your life _as if_ you were alive in the real world. 3. This may not be as exciting as the thought of being in the afterlife or in a simulation. Maybe that was part of the reason why you were compelled towards such beliefs in the first place... If it's so, then consider finding joy in other things which do not involve unprovable beliefs that leave you confused about the matters of reality.