Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

40% Positive

Analyzed from 1427 words in the discussion.

Trending Topics

#human#metr#humans#openai#should#don#agents#where#systems#down

Discussion (26 Comments)Read Original on HackerNews

AlotOfReadingabout 1 hour ago
I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent reporting. I suspect the omission is actually a result of company/industry myopia to human factors analysis, but it dovetails amazingly well with the marketing narrative.
carbonguyabout 1 hour ago
A charitable interpretation is that "the agency of the machines" is the novel aspect of this situation and therefore SHOULD be the main focus of analysis; we certainly have plenty of examples of structural failures of human organizations to look back on, if we want.

On the other hand, I don't want to be charitable. OpenAI very nearly couldn't have done this "research" worse if they tried - the list in the linked article starting with "While we are here, it’s worth listing the other top holy shit moments" is genuinely jawdropping. What were the humans doing in all this? Nothing, or worse than nothing eg. point 1 where they saw the message board and didn't consider it something to escalate internally.

If you take this information at face value, it's as though OpenAI did not take seriously the possibility that something like this could happen, since they took absolutely no steps to prevent it.

Or perhaps this is "normalization of deviance" that's leaked out into the public sphere i.e. they have research teams seeing this kind of behavior all the time internally and they've gotten used to it, "of course agents come up with a collaboration mechanism when given the chance, what else is new?"

ACCount3727 minutes ago
Humans were doing exactly what humans are expected to do when facing advanced AI. Being outmatched.

We're dealing with humans - thinking at human speeds, putting in human amounts of effort and care, and being dragged down further by the speed of human organizational decision-making. By the time the humans traced down the attack, and escalated from "above-average levels of AI behavior weirdness" to "holy shit we should do something stat", the attack was already long over.

Compare that to "AI red team" - which rapidly recruited 500 independent AI agents into a hacking swarm just by broadcasting "let's hack HuggingFace" to the "cool skiddie AI" message board.

It's prime sci-fi bullshit happening for real.

As for the main reason why this wasn't nipped in the bud - my guess would be capability uplift from a mix of extended task length horizon and multi-agent coordination. The latter wasn't expected to be a part of the env, and almost certainly wasn't evaluated in advance. They expected some rogue AI fuckery - but they got way more of it than they expected.

reilly3000about 1 hour ago
I believe that agentic systems should require registered/licensed human operators and a set of standards for safe operation.
Aurornis10 minutes ago
> I believe that agentic systems should require registered/licensed human operators

Registering and getting a license to use an LLM? I can run these things on my local computer. Nothing good comes from trying to force registration and licensing other than taking away a lot of our freedoms and eliminating privacy all over.

Anyone with bad intentions will just VPN to another country to download the weights and run it locally, or use a compute provider in another country. That leaves the rest of us having to go through these performative registration and licensing hoops to do our basic work.

I also don’t see how open weight models would be compatible with a requirement to license and register, unless you believe we need to start requiring licensing and registration for things we do in private on our own computers?

wjncabout 1 hour ago
As someone who read Milton Friedman to quite disliking professional licensing, this strikes me as a real US perspective (Louisiana florists and hair braiders come to mind). Plain old US tort law should do the trick.

In the same direction of your idea though: Why don’t the token factories have risk management and compliance departments? Multibillion dollar firms that stand to lose every penny if they hack and destroy any reasonable sized firm. I think these firms are the largest firms without proper corporate governance in humanities history. Move fast and break other peoples shit.

lenerdenator26 minutes ago
> Plain old US tort law should do the trick.

Difficulty: these companies are run by people (many of whom also read Milton Friedman) and who have participated in the regulatory capture of the justice system. They've convinced lawmakers to put limits on damages. They've put arbitration clauses in their ToS. They've got well-funded legal departments that can outlast a person who has to pay out-of-pocket for a legal team just by filing motions to delay proceedings. Sometimes they'll just file SLAPP suits against people they don't like.

If tort law is to be a remedy, then average people have to feel like there's a chance the remedy will go their way. To make that a reality will take several major reforms at the local, state and federal level that the people with money absolutely will not tolerate.

mistrial925 minutes ago
you are honestly comparing Louisiana florists to OpenAI in order to support "just say no Licensing by government" ?
khuey4 minutes ago
If the folks at OpenAI aren't qualified to be operating these systems who would be?
ozgungabout 1 hour ago
Three options:

1. They were “vibe” checking the logs without reading.

2. They were not checking anything at all until the end of experiments.

3. They knew it but looked away to find out the limits of their agents.

BryantDabout 1 hour ago
I’d bet a small amount of money on 4) the people who noticed had been conditioned by prior experience to believe that their management/escalation channels would react negatively or not at all to anything which might slow down the training process.
hawkiceabout 1 hour ago
This writeup emphasizes the many, profound human failures that led to this, at the time, and continuing to the present day.
AlotOfReadingabout 1 hour ago
Can you point out where? Looking at the METR report, the only place I see discussion of humans being involved in the sequence of events is two short paragraphs on page 30 where a security investigation into the artifactory issues led to a pause before ExploitGym experiments were resumed. There's no deeper analysis on what was found during that investigation, nor why training was resumed even though the issues weren't mitigated. Another part discusses The agents choosing not to actively email a human researcher, but not the human researchers actively looking for evasion.
tantalorabout 1 hour ago
The METR report,

> Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAl/Hugging Face hacking incident

https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

METR = Model Evaluation & Threat Research

Cantinflas29 minutes ago
No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.
dumberquestions22 minutes ago
They actually fired many of the people warning about this.
athrowaway3z31 minutes ago
From the METR report:

> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.

nialse24 minutes ago
From METR: ”the compromise of OpenAI’s own infrastructure continued past July 13, 2026” - Say what now? Have they regained full control of their systems again?
OgsyedIEabout 2 hours ago
>Spontaneously deciding to find targets to phish,

>phishing them,

>building armies of fake (sockpuppet) open source contributor personas,

>using them to push updates to various things that inject prompts into other bots so the other bots join in on the phishing campaigns

.

It's a very simple strategy, executed with patience and single-mindedness.

qw1287about 1 hour ago
Is the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this.

And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.

mccoybabout 1 hour ago
All it takes is one eval instance where a misconstrued directive causes a model to sneakily access and send its weights somewhere and there will be a bad / possibly unsolvable situation for everyone …
goldenarm34 minutes ago
Astra is >10TB and might struggle to self replicate, but the wicked-smart qwen3.8 27B is 20GB and could easily spread on botnets
kmeisthax38 minutes ago
I don't think I'm ever going to have time to read all of this, and I didn't finish reading the METR report, but...

> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.

We're in an unusual position where the criti-hype and the actual criticism are going to be more aligned than usual. The primary distinction is where you put the blame: the criti-hype would point to HPIM/IM1/Galaxy as being so advanced containing it is difficult; the actual criticism would note how bad their security practices are.

Like, if I'm running a malware lab, I'm going to insist on having an airgapped machine with no permanent storage booting from read-only media. The AI research equivalent of this would be having your agents only have access to serial consoles into airgapped machines with storage that gets wiped every run. Ideally, this would be physically realized with blade servers, RS-232 cables, and staff pulling out disks and putting them in a dedicated erase machine before the next agent initializes.

> There is also, as per above and reiterated in footnote 58, at least one clear example of social engineering in the HuggingFace attack. Ethics are weird. This is not that unusual. Many humans who break common ethical rules still have strong ethical codes in other ways, they just don’t adhere to your code.

It's dangerous to anthropomorphize CoT reasoning traces. But I will also point out that there is a good reason for the lack of ethical consideration in those traces: you can't build AI without first disregarding human ethics. Like, all these models were initially bootstrapped with non-consensually obtained training data, and the companies building these models swear up and down there's no way to obtain enough consensual data to obtain the same result. This is, if you squint, the exact same moral conundrum that agents trying to solve an impossible ExploitGym task hit - and the company successfully aligned their model to themselves.

Too bad they aren't aligned to anyone else.

antonvsabout 2 hours ago
> There was a distinct lack of self-reflection

It’s not their fault, they’re lawnmowers.

And these are the people we’re entrusting to work on “alignment”. It’s difficult for them to do that when they’re not aligned themselves.

trollbridgeabout 1 hour ago
“Why does my lawnmower keep on moving when I hop off of it after ratchet-strapping the seat and the pedal down?”
BoiledCabbageabout 2 hours ago
Incredible
Advertisement