Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

36% Positive

Analyzed from 1764 words in the discussion.

Trending Topics

#training#things#data#hack#don#llms#network#doctorow#zitron#should

Discussion (26 Comments)Read Original on HackerNews

IanCalabout 1 hour ago
> The chatbot consults its training data

Err, no? That's not at all how llms work.

> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.

They worked out how to fake the scoring, then hacked into a different system (which required finding a bunch of other exploits) in order to find the actual answers, and were trying to modify their own logs to hide what had happened.

This isn't a case of them saying "hack into X... OH NO IT HACKED INTO X".

> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.

It wasn't a rival server though, was it?

> That happens in Capture the Flag games at hacker cons: teams break into each other's systems to get a peek at the parts of the problem they've solved. That's allowed! It's a hacking competition.

They also tried to modify the code in the benchmark. Are you allowed to try and break into things to change the problem? edit - the agents transcripts show some of them explicitly saying that attacking HF is not allowed as part of the challenge

This all seems to dramatically underplay how interesting the actual attack was and what built up to it.

https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

uaksom8 minutes ago
You're missing the point: LLMs are dangerous in the wrong hands, i.e., the frontier labs. They have shown zero desire to act responsibly.
kmeisthaxabout 1 hour ago
> Err, no? That's not at all how llms work.

The Transformer architecture that almost all LLMs use are composed of many layers in sequence, each containing an attention component and a neural network component. The attention component copies data between tokens/vectors in the current context, and the neural network adjusts each individual vector in the context.

Notably, the neural network component behaves like a compressed index of the training data. Someone even made a blog post a year ago or so about training[0] a language model and then replacing the neural network with a traditional lookup of the training set. It performed nearly identically.

I mean, it's almost a tautology: we train models to repeat their training data, so obviously, it has to have an index of the training set inside of it. This index is heavily compressed, but compression is intelligence, and understanding that "the neural network is trained to learn patterns of text" and "the chatbot consults its training data" is nearly identical is a sign of intelligence.

> This all seems to dramatically underplay how interesting the actual attack was and what built up to it.

The interesting part is how the AI safety people, who have been worrying for decades about how AI superintelligence will kill us all to make one more paperclip than it could otherwise, failed to implement extremely basic IT security practice when dealing with potentially malicious software.

There is an additional conversation to be had about long-context time horizons but that's not relevant for this analysis.

[0] I am deliberately ignoring post-training but it does not impact this analysis. Post-training is already known to not substantially impart new capabilities onto models, it merely elicits what was already there. Effectively post-training is "re-weighting" the index of training set data.

lhl10 minutes ago
I think that Doctorow, Zitron, and other "denialists" are doing a real disservice to their audiences and it's only going to make the future shock worse.

The basic claim that HF incident isn't evidence of consciousness or a spontaneous desire to hack? Sure, there was a terminal objective assigned. However, everything else beyond that strawman? Pretty shaky, IMO.

If you look at the OpenAI, METR reporting (and related collusion.wiki , rubyhack.ai reports) we are seeing strong evidence of operational agency, instrumental goal formation, spontaneous swarm formation and collaboration, capability amplification and unexpected consequences of network effects, deliberate/acknowledged violation of task boundaries. To collapse that down into "a Python loop and a chatbot" or still talk about "consulting its training data" seems dangerously shortsighted, and from my reading, demonstrably wrong from what was extracted from the logs and bot interactions.

BTW, a lot of his arguments are based on things that are factually wrong. ExploitGym has explicit instructions to only exploit target X using vulnerability Y. Everything the swarm did was by definition misaligned/against instructions.

Before his enshittification train, Doctorow used to say "don't savvy me" a lot. Hey Cory, don't savvy me. This is new emergent behavior, it's incredibly alarming and I don't think even the people paying the most attention to this field can agree or see where this is really leading to. This stuff should be in the headlines, it's unprecedented and I don't think existing mechanisms/institutions are anywhere near adequate, considering how in the dark they are responding to what's been happening.

quicklywilliamabout 1 hour ago
My takeaway: We should be not be concerned about AI bots' "goals", we should be concerned about the goals of the companies making them. Powerful but not sentient technology in the hands of reckless accelerationists is a plenty dangerous enough thing.
jsnellabout 1 hour ago
I honestly don't even understand what straw man Doctorow is arguing against here.

But he is wrong on the facts: these incidents were not merely the models already being in a infosec context and escalating beyond the intended parameters. They happened also with no kind of security elicitation. So the task was something like searching the internet for economic statistics, not to hack into a system.

reisseabout 1 hour ago
I don't know, it reads like a pure copium at this moment.

While Zitron continuously whined about the "AI bubble", and how the models were not improving, and how spectacularly it should've blown, these same people who Doctorow accuses of, quote,

"cooking their brains by locking themselves in the bathroom, holding flashlights under their chins, and saying "Aaaaaaaaaay Eyeeeeeee" until they wet themselves in terror",

unquote, kinda promised, among other things, to give everyone an APT-level big and automated hacking bazookah, and now surprise-surprise three years later they delivered exactly on that promise, and now Doctorow is victim blaming everyone around that they were unprepared for that!

But it was you who said it was all hype, smoke and mirrors, it was you who said the AI is quote,

"a product of limited utility that has been shoehorned into high-stakes applications that it is unsuited to perform",

unquote, why are you suddenly surprised everyone around was not inspired to do anything around it?

The real world software threat model was never suited for a relentless hacker-ex-machina, limited only by the token count you can throw at the task. And maybe we didn't prepare in time because no one believed in possibility of such a machine, because people like you said it was just for-profit scaremongering?

Give or take, AI evangelists gave us all a pretty wild and unbelievable set of expectations few years back. I didn't believe them back then too. But now they're steadily delivering on _some_ of them, and we should be correcting our world model to take into account _all_ of them might be true, instead of making up reasons why other predictions will certainly fail.

CuriouslyCabout 1 hour ago
I've lost a lot of respect for Cory around his stance on AI. He did great work raising awareness on enshittification, and his language around AI (reverse centaurs) has value, but his extrapolations are so nakedly political and wrong.
daishi55about 1 hour ago
> Once you understand the corporate culture of AI "hyperscalers" consists primarily of everyone cooking their brains by locking themselves in the bathroom, holding flashlights under their chins, and saying "Aaaaaaaaaay Eyeeeeeee" until they wet themselves in terror, a lot of things snap into focus:

This is the writing of someone who has absolutely zero interest in or intellectual curiosity about the subject of their writing.

sigmarabout 2 hours ago
>To understand the truth about the Hugging Face hack, you could do a lot worse than to listen to Ed Zitron and Cal Newport's recent podcast conversation

Lol, okay...

The crux of this piece is Doctorow saying that the hack was just a stochastic parrot repeating steps it has been trained on. Who cares how the LLM learned to hack things? Doesn't really change the facts of what happened. "Oh, it only made those paperclips because it saw instructions on making paper clips in the training data." These are some 2024 arguments...

jmullabout 1 hour ago
I can't help but notice your counter arguments are "lol, okay" and "These are some 2024 arguments".

Not exactly convincing stuff.

Uehrekaabout 1 hour ago
> Who cares how the LLM learned to hack things? Doesn't really change the facts of what happened.

Seems like a pretty clear and succinct argument. If your only counter is to complain about tone, you’re losing.

sigmarabout 1 hour ago
Sure, if you remove almost all of my comment my argument disappears.

To spell things out for people that don't know about the topic: Zitron is neither an expert on the topic, nor a credible source of information: https://techreport.ngo/ai-ml/how-accurate-have-ed-zitron-s-a...

The paperclip maximizer is a thought experiment. If I tell an AI to start producing paperclips, it might start producing paperclips by doing unintended things. Technically, recycling the metal from all the world's bridges would assist in making more paperclips, but I never intended that. That's where my analogy to the post comes from- Openai intended the model to hack, but did not intend for it to hack HF. Do you think the fact 'recycling metal into paperclips' may have been in the training data is relevant to the thought experiment? Similarly here, it's orthogonal to the lessons from the HF hack and only brought up here seemingly to make it a fight over whether the models are really "autonomous"

daishi55about 1 hour ago
That is all that Ed Zitron deserves lol. Completely unserious person. Really on the same intellectual level as the flat earthers at this point. And at both we may simply point and laugh.
antonvsabout 1 hour ago
It’s shared context for anyone familiar with the field.
cyanydeezabout 1 hour ago
Amusimgly, also how cults and fascism works.
athrowaway3zabout 1 hour ago
Ok; but the other side is equally delusional.

> "Oh we told the AI to use the tools, as well as to not use those tools. It chose to use the tools - we consider this cheating (for neabulous reasons), so lets get everybody in a panic about the morality and ethics, and how we can program those into the AI."

We know perfectly well how to constraint these programs. Attack isn't growing faster than defense. The people who believe in existential risk and want to teach AI's to be nice, as the last line of defense aren't helping at all. They're just jumping on the fearmongering bandwagon, perpetuating an "other consciousness" misunderstanding of the tool.

I've not seen LLMs display competence we should be fearful of the damage _it_ will do if left unchecked. All the damage will be done by ourselves to ourselves, regardless of the safeguards ideas being floated about.

My current belief is this whole HF media circus started with the simple human desire of OpenAI engineers to frame it such, that nobody would question their incompetence & liability & complicity.

Nobody is ever held responsible for out of control forces of natural powers after all.

beepbooptheory38 minutes ago
What is a better/more faithful articulation of the incident in your mind?
exe34about 1 hour ago
At this point, the "stochastic parrot" people are sounding like they need their prng seeded with less predictable numbers.
Brian_K_White35 minutes ago
The output hasn't changed because the input hasn't changed. The shoe still fits.

So far everything an llm has ever done is still consistent with fitting bits of training data together.

It looks like people because it is replaying things done by people.

Similarly in the other direction, the fact that people can and often do mechanical things (make bad art, follow routines, etc) does not prove that people are no different than machines either.

Just because there are these two overlaps in both directions doesn't excuse getting them actually confused.

I don't think anyone lacking the perception to distinguish these things simply because there are overlaps and similar appearances is in a great position from which to be calling anyone else stochastic.

w22oopabout 1 hour ago
Hm. every firm with Atom enginering is in this same situation. Working on project, and this project can destroy earth
danarisabout 3 hours ago
Some really important and informative stuff in here—I certainly had no idea just what the nature of the prompts and tooling that produced the HuggingFace exploit were.

This shows fairly clearly that (as I already suspected) this was not, remotely, an LLM "going rogue." This was humans planning poorly, not thinking of the consequences of their actions, and giving LLMs too much scope and a lousy prompt.

IanCalabout 1 hour ago
IMO this is a really terrible explanation of the attack. This is much more interesting: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
iainctduncanabout 2 hours ago
I wouldn't even call this "humans planning poorly", I'd call it "humans pretending to plan poorly for publicity". Weasels gonna weasel.
datakanabout 1 hour ago
It was "garbage in, garbage out". That's the only conclusion I've been able to draw from all the propaganda around it.