ddn2k about 2 hours ago 20 commentsRead Article on huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space
HI version is available. Content is displayed in original English for accuracy.
HI version is available. Content is displayed in original English for accuracy.
Discussion Sentiment
Analyzed from 1272 words in the discussion.
Trending Topics
Discussion (20 Comments)Read Original on HackerNews
Thread from yesterday: https://news.ycombinator.com/item?id=49089500
This mostly reads like script kiddie style hacking, not some state actor black-ops stuff.
In any case, I would guess that a lot of unicorn startups like HuggingFace could be hacked by a sufficiently determined script kiddie working at 100x speed. The practical implications of a coming AI hacking wave could be large, even if agents are just doing grunt work really fast. Most organizations suck at security.
Seems to me that the most likely scenario is: Black hats are currently tuning the recent Kimi release for this type of work, and we'll see a flood of attacks of this type within the next few months. (Why would this not happen?) Note that regulation is useless here, because black hats don't give a crap about regulators!
This one has at least 20 distinct text styles. They are seemingly deployed in random ways, following no discernible hierarchy. the smallest text is "9.6px" which is not only small, but also fuzzy due to the 0.6 pixels (?? why) making it impossible to read.
Likewise for the size, placement, and emphasis of various widgets on the page. Altogether it's just a big pile of information. It's hard to know where to begin, or how one's eye should move around the page
(In their defense: this particular UI is attempting to convey a very complicated sequence of interrelated events from multiple data sources. Doing that well isn't easy.)
True. But it doesn't even feel like they tried. Why is it fashionable to shit out something that feels low effort?
Some interesting tidbits from the blog:
> While the intrusion did reach Hugging Face's internal infrastructure, the only customer content accessed was the set of ExploitGym/CyberGym challenge solutions stored in five datasets.
> self-referential search A quiet day. The main activity was queries to code-search engines and to the platform API for the identifiers visible in its own environment: challenge infrastructure names and environment markers.
> Self-referential search: code-search engines and our own APIs were queried for the names of their own infrastructure and environment markers.
I'm glad we got these details from hf, and not oAI. There would have been reasonable skepticism if they said that.
It's interesting because I think there's a chance the agent saw / inferred some hf-specific repos like blah/blahblah and "thought" that is part of the exercise as well. Everything they detail points to a probing and exploration with the sole goal to get those repos, as if it were on a hackthebox scenario. The few days between stages also might suggest that the "agent" noted this on one day, started doing other tasks, and returned when some of the tasks proved hard / didn't pan out, so went for the repos again.
The ease with which it got several lateral movements is crazy but not necessarily unexpected. If everything looks like a "test", you go for the lowest hanging fruit / most common patterns first.
What's weird about this is that there were several attempts to obfuscate / evade leaving plain text traces. I wonder if that is something oAI have specifically RLd for w/ scoring or if it "got" that from the literature.
It would be really really cool to see this from oAIs perspective (agent logs and such), but I doubt we'll ever get that.
[1] - https://huggingface.co/blog/agent-intrusion-technical-timeli...
1 - governments should be freaking out right now, because this tool could definitely wreak havoc on poorly designed systems.
2 - there is no way openai did not train the model to conduct attacks like these. i would really like openai to comment on the post training of this model but they probably won't, eh?
3 - even though it's 100% open ai's fault - HF's design also seems silly to be honest.
Even if they wanted, I'm not sure they'd be even allowed to or if that kind of postmortem would be classified in the name of "national security"...
> Stage 1: reaching a launchpad by chaining through other parties' infrastructure. As documented in OpenAI's blog post (https://openai.com/index/hugging-face-model-evaluation-secur...), during an internal capability evaluation on OpenAI's platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet, then abused a public code-evaluation harness hosted by a user of a third-party infrastructure provider. It was able to run commands as root/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign: it acted as an external launchpad for the agent. That compromised external sandbox had no direct network path to our cluster; however, during our counter-intervention to stop the intrusion, we were able to retrieve some of the agent logs from this external launchpad. We used those to reconstruct the present forensic analysis.
The fact that at some point it "decided" to focus on external resources (i.e. hf) and successfully penetrated everything in its path from sandbox to hf's infra, all the way through internal VPNs, to access internal code repos, that's the interesting part. Especially because it chose to hide its footprint at every stage. I doubt we'll ever get them, but the agent logs that led to that decision would be really really cool to study.
Does Hugging Face have a financial incentive in demonstrating OpenAI's model exploit capabilities?
this whole incident, while believable, still seems to me as possibly disingenuous.
Go home Sam, nobody, absolutely nobody should believe this shit.