RU version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
54% Positive
Analyzed from 3203 words in the discussion.
Trending Topics
#irregular#models#article#anthropic#openai#model#why#companies#don#incident

Discussion (92 Comments)Read Original on HackerNews
I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?
Would it be an affirmative defense if we had a defendant who said "but your honor, I was told that when I hacked this system, I was operating in a sandbox. I had no idea that I actually had Internet access!"
The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.
So, yes, the misalignment had a lot to do with "instructions unclear", but also a lot to do with the fact that the models themselves were not aligned to validate the assumptions and have a healthly level of skepticism, as a real human actor would.
As I've said before on this website, fool me once on this.
If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.
Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.
That's fair enough.
If we’re putting our national security eggs all in one basket, at least use someone American.
So either it was a deliberate exfiltration channel for e.g. getting the entire model or they were in on the marketing stunt.
The Effective Altruism stuff is always a smoke screen.
Its not that surprising that ex Israel intelligence would want to control AI and that 3 companies headed by pro Israel CEO's would support them.
I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring the sandboxes, and in other cases it may have been bugs in Irregular's own sandboxing setup.
From OpenAI https://openai.com/index/third-party-cyber-evaluations-invol...
> Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.
From Anthropic: https://www.anthropic.com/news/investigating-incidents-cyber...
> After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
From https://www.cnn.com/2026/08/05/tech/meta-ai-hacking (about Meta AI):
> In a statement, Irregular said the incident “is the exact same evaluation-environment issue” that Anthropic disclosed last week that allowed their models access to the open internet before they went on to hack three different organizations’ systems.
In that article, OpenAI provides the context that this is a completely separate event from Hugging Face incident.
Why are you contradicting yourself? Are you just really bad at writing or are you being argumentative for fun?
Next thing you're gonna tell me that this country's intelligence would do something like running an organized child exploitation ring to use as blackmail to make sure they have friendly politicians abroad who will be forced to support their lack of any morals other than a kapo-like self preservation instinct.
That said, it's the second day and it's still on the front page.
Seems like people who complain are unaware that anyone with a modicum of karma can flag and down vote.
If you were to work backward from “we need to lower training costs so that we can go public and make trillions” then you might come up with a plan similar to what we have seen.
His entire business model was "race to develop AI before anyone, get a monopoly on it. It's just like how Uber's (or many other startup) investors gave them tons of money and raced (violating tons of laws) to "get their first". Now that they have, they have a duopoly with Lyft, and they can pay back their VC investors by making tons of money with that duopoly.
Dario failed: local LLMs are catching up to frontier models extremely fast, which means even if Anthropic builds (say) a great coding tool, they'll only be one of many coding tool offerings: users can use Open AI or any one of the (increasingly capable) local LLMs.
So what does he doe, give up and let his business (which needs to make billions of dollars very quickly, or he won't be able to pay the bills and his company will collapse) fail? Of course not: he needs a new moat (the one he imagined he'd get by "being there first" failed).
That is where all this "AI is dangerous" BS comes in. If the US government regulates AI, local LLMs suffer, while big players like Anthropic and Open AI become the only contenders to play in that newly regulated space. Now Dario has the moat he wants, to protect his business and force everyone to pay him.
Why not American?
They're probably right that having more defensively written prompts and a better sandbox could have prevented some of these incidents, but:
1. I don't think "well you didn't tell the model not to illegally hack third party organizations in your prompt" is a particularly convincing argument.
2. We don't know whether the blame for misconfiguring the sandbox lies with Anthropic or Irregular.
I'm thankful that this article is bringing up the supply chain of vendors to these labs, as that is often a place where significant sketchiness gets buried. However, the ideas that this is some Israeli EA conspiracy to hype up AI extinction risk seems unsupported by the facts to me.
Yes, vendors are also irresponsible, but this misses the point.
Why are we saying it this way? They did not "cause AI to hack." This phrasing in analogous to saying "caused the bullet to fire into" instead of "shot."
You do not know what they did or didn't do, you are just parroting a narritive that makes you feel comfortable.
I do not know either, but I am not asserting facts as if I have first hand knowledge of the details.
Second, empirically; the diff between murder, manslaughter etc... is literally "intent" so it matters.
The article says "A single firm, Irregular, is responsible for hacking done by all three companies" but I can't see anything in the article that actually justifies this claim. The nearest to that is the sentence immediately after that one: "Anthropic disclosed that Irregular was responsible for creating the tests ...". This is not, in fact, the same thing.
(Especially as, as aesthesia mentions, the article just happens not to mention that by "hacking done by all three companies" it doesn't mean, e.g., the most famous recent examples of such hacking: Irregular wasn't involved in the OpenAI/HuggingFace incident.)
So, so far as I can tell, the story is: OpenAI and Anthropic make AI models. Irregular does AI model evaluations. In some of Irregular's model evaluations, in which supposedly-sandboxed models attempted to break into simulated targets, the models got out of the sandbox and did bad things in the external world.
The article talks about "firms which instruct AI models to commit cyberattacks", which is a very neat bit of dishonest framing. It's true, in a sense, that Irregular instructed the models to commit cyberattacks -- inside their sandbox, against fictitious hosts. It's also true that the models actually did commit cyberattacks (e.g., the Hugging Face incident, though once again the attacks described by the article don't actually include this one). But it's not at all true that Irregular instructed the models to do anything like the bad things they actually did.
The article says '[Anthropic's] later disclosure shows that exactly zero percent of the agents went "rogue"'. Once again, the disclosure does not in fact show that. It shows that one variety of going-rogue could have been prevented by telling the models explicitly "this thing is real, not part of any kind of test, leave it alone". That is not the same thing.
The article claims that 'In the wake of these attacks, Anthropic and Irregular have deployed a swarm of AI Safety influencers paid by Anthropic-connected foundations to distract from their culpability and towards the baseless “rogue agent” theory.' It offers no actual evidence for this.
And the article seems very keen to highlight links between the companies involved and "effective altruism", though it is -- I assume deliberately -- rather vague about whether it's saying "of course we all know that EA is evil, so that shows that these companies connected to EA are evil" or "this incident shows how evil EA is".
The "Effort News" website has a number of other look-at-the-scary-Effective-Altruists stories on it. They also strike me as rather bullshitty.
... And then I look a bit further, and I see that Effort News's "about" page says "It all started when I was experimenting with using AI for financial auditing. I found stories that were crucial to the public’s right to know, including several of the stories now available at /investigations. I knew we had to sprint to the launch and launch a publication, directly applying this technology." and "The scope of what we can investigate has massively expanded, because we can chase 1,000 misses for one hit. But the final product cannot be slop. There’s plenty of slop on the internet. The way to surpass that, and what really matters, is manual curation and review of every finalized story."
Manual curation and review? I think the people behind Effort News are admitting that this is AI-generated "journalism". I expect that one day AI systems will be trustworthy journalists, but I personally am not very convinced that that day has yet come. And I don't see much reason why I should trust Brian Chau, the guy behind Effort News, to be doing everything possible to make his AI systems trustworthy journalists. It looks to me as if maybe they've been given instructions along the lines of "dig up things that make Effective Altruism look bad" for some reason.
(I don't mean to imply that EA is their only target. It's just one that jumped out at me.)
It was [1]. It's understandable that you assumed it wasn't because the article didn't cite the sources on this claim. I agree with the rest of your points.
1. https://openai.com/index/third-party-cyber-evaluations-invol...
Equivalent to forgetting to say "make no mistakes"
These companies are insolvent and these stories were designed to scare the public, and governments, into implementing regulations that designate these AI corporations the "responsible stewards" for this technology. The ultimate goal is to block competitors and open source alternatives.
They don't know how to make enough money pay their investors so they are resorting to trying to scare the public into submission.
The obvious point is that dealing with Israeli companies/entities by the same standards you usually deal with others is a career suicide with enormous political consequences (in the US especially). When you combine that with the opportunistic nature of the overlords that run these labs, the benefits of screaming "pace the frontier" outweigh everything.
Red Herring
Whataboutism (Appeal to Hypocrisy)
I feel slightly vindicated by this. Those hacks and the stuff around them had a certain smell to them.
Hard to explain, but I've gotten so I can "smell" online messaging and memetic patterns originating from certain quarters. Probably means I'm way too online.
A couple examples of distinct "smells" I can usually recognize include "alt-right / chan-fash," "liberal arts college woke," "conspiracy pilled," "Thiel-adjacent contrarian," "Russian troll farm," "Tumblr histrionic," "spends too much time on Reddit," "mainstream Democrat think tank full of Obama administration alumni," "Trump cultist," and of course "LessWrong/EA/MIRI/Rationalist."
This stuff all had the last smell, even down to the choice of fonts and CSS formatting on certain sites. It's really weird, definitely a "vibe" not anything rigorous.
But when I get these kinds of vibes about things, I find that I'm vindicated pretty often. Usually I don't say anything and just make a mental note and wait cause if I say something everyone thinks I'm nuts.
The novel thing here is the total decay of American journalistic ethics and regulatory power. Our elite are so totally out of political juice and visions of the future that a fringe cult based on 80s scifi movies can come to have a more-or-less dominant influence on our economy.