Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
50% Positive
Analyzed from 21064 words in the discussion.
Trending Topics
#agents#more#why#models#human#don#humans#problem#llms#training
Discussion Sentiment
Analyzed from 21064 words in the discussion.
Trending Topics
Discussion (653 Comments)Read Original on HackerNews
LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.
This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".
We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
"Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action, passivity.
Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.
OpenAI didn’t ‘let’ these bots do this any more than someone ‘let’ Claude Code make them a website.
It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild.
But some automation is different. The most prominent example before AI would be car navigation systems, where the entire idea is that that you give it a destination and it figures out the exact actions to get there on its own.
Except even there, the actual driver would still have been you - giving you a chance to vet and deny every turn the system proposed.
AI agents are sort of like that - most of the value they provide is in the ability to turn high-level goals ("write me a traffic control system for my model railway") into low-level actions and also do so interactively.
The new thing is that the "driver" has much less oversight here where the agent wants to go, and is sometimes removed completely. That part is clearly be an active decision by AI labs.
The other thing is that the labs seem increasingly to steer their training towards behavior that make events like this one more likely, e.g. that agents should never "give up" when faced with a seemingly impossible task, but instead should keep trying and think of increasingly outlandish ways to solve the task. To me, that seems pretty much a recipe to get incidents like this.
I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!
I’ve said this before in another thread and people went absolute apeshit saying it is an unreasonable expectation and that AIs absolutely REQUIRE this anthropomorphic human-like speech pattern to function correctly.
I cannot overstate how deeply wrong they are.
But into today's AI climate it's becoming increasingly difficult to figure out who is shilling, who is being assinine and who actually believes AI could do these things without clear human instruction and enabling.
OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used.
If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn't speak to the intelligence inherent to the system.
Not that the techniques used by the LLMs in the actual incident weren't unexpectedly sophisticated, but the outputs of each and every one of these processes could've been read at any time during the run. They just weren't.
If the robots obtain sovereign nationhood, and are able to self-sustain, then autonomous robot decides for itself will be a valid argument.
I got fired for some reason.
Recognizing that the models are acting with intent does not somehow absolve OpenAI from their felony hacking. We have not granted them personhood.
The car doesn't have agency, it's doing what it naturally does. LLMs are the same, they're working as designed.
But I don't understand the point of splitting hairs. You are always responsible for the actions of your devices, tools, machinery, software, employees, whatever.
Trying to blame AI for one's own stupidity must be aggressively pushed back on at all times.
If your buddy leaves his car parked at the top of a hill without the parking brake on and it rolls down the hill and side-swipes a bunch of vehicles and narrowly misses an elderly person walking by with a cane someone could easily say:
"Dude wtf is wrong with you, you left your car parked on the top of a hill with no brake and let it roll into traffic"
The phrasing doesn't absolve the offender of their negligent behaviour and the consequences of it.
The only thing thing does is the lack of action from regulators and society writ large.
Our lack of action is what allows people like Sam Altman and Dario and the irresponsible people who choose to work for them to be continue to be negligent.
So if the guardrails suck, or they're left off for research purposes, bad things can happen.
A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance.
I have never had an issue with agents doing something they shouldn't because I observe them, and I leave the vendor guardrails in place.
I can understand agents coordinating in unsupervised scenarios: I would see it as an aspect of intelligence. We ourselves build up knowledge by reusing what someone learned before us.
Einstein, other greats, always stand on the shoulders of other forgotten giants. Other discoveries by other people taken as fact, so that we can build some new ideas on top.
Agents swarming amd sharing solutions to problems is more efficient, the same way it's been efficient for us.
Reaching out for help in this way is like probing the air in the dark with your hand: sometimes your hand hits something (another agents solution to a problem) and so you can use the info to adjust your own motion to get to where you need to be faster than if you just run full speed into everything.
Electricity follows all paths, not just the one with least resistance.
Except in a handful of limited cases, eg. medical and aviation.
It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability
If you want to test military missiles, you do it in the f'ing desert, not from New Jersey.
You want to run ai without guardrails, do it in an airgapped system or be held accountable.
0: https://en.wikipedia.org/wiki/Marcus_Hutchins
1: https://en.wikipedia.org/wiki/Tornado_Cash
If somebody else runs the software, then they are.
It's both, isn't it? For example, in very early days of agentic coding, I once had a rule saying "don't read or write any file outside your current working directory." Then AI just wrote a bash script and access those files anyway. Did I 'let' it do it? Technically yes. Did I know how to set up a sandboxed VM? Also yes. But how were I supposed to know that it could and would do that as someone new to this tool?
It was a genuine eye-opening experience to see AI just do things in ways I were too complacent to expect. I kinda expect the SOTA LLMs would find a way to escape my VM and access files on the host system (haven't tried it though).
Both?
The AI companies act irresponsible, but it is still very interesting how those agents can behave?
LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.
Whether they have a soul or consciousness or feelings doesn't matter here, because this is what they did - and this is very dangerous behavior. Especially with all the irresponsible people in power right now all over the world.
I think a little bit of humility for the capability of these machines is warranted at this point.
> immediate anthropomorphisation
Ok, why don’t you try?
You could hire a legion of people for cheap to commit these crimes, but if its bots, suddenly its unpredictable and just one big whoopsie and therefore perfectly ok to do?
If a human starts pentesting a site its a crime, but if a bot does it its an AGI frontier doomsday scenario and that automatically pushes the consequences off their table?
Whats going to happen next? Are we gonna have robots that happen to physically break into banks to rob them for some reason and the company making them isnt responsible just because?
I've been saying since the start that AI is a tool that a human is using and should be treated as such. They should carry the responsibilities and the benefits. That way our stance would be consistent.
Maybe some analogy could be with children - as a parent, you are responsible for their misbehavior, but their achievements are their, not your?
Nuts to that. We should be interested in why things happen, not just finding scapegoats.
If I ran Metasploit against HF and RubyGems because I “accidentally” misconfigured my lab sandbox, there’s a good chance I’d be prosecuted.
I don’t think LLMs vs Metasploit being different software changes the law.
The problem IMO is the executive. The DOJ is declining to take any action against frontier companies (aside from possibly Anthropic) as the stance of the admin is that the companies are "critical for national security". For example, see the DOJ's request to dismiss the NAACP datacenters lawsuit against xAI: https://www.utilitydive.com/news/doj-intervenes-xai-data-cen...
What does "running a Wuhan" mean?
How would that work out if these were self-hosted open weight models?
Of course in cases of negligence a tool maker could also be held partially liable. That’s a matter courts can decide. The main point is we shouldn’t jump to making special laws around the development of LLMs. The starting place should be enforcement of existing liability laws. New laws take time and will be heavily influenced by AI companies seeking a regulatory moat for their business. Moreover, it is a distraction from the illicit behavior that is already going unchecked.
Isn't that what the goal is, though? we're coding them to close the delta between what currently exists and some nebulous end-state - to me, that sounds like a formal definition of 'desire'.
AI is a tool, and it won't be long before the damage caused by its improper use affects real human beings. These were warning shots.
The most absurd part is that everyone agrees, governments and AI companies included, that the scale of the potential damage and the long-lasting effects of losing control of AI should not be underestimated. Yet, at the same time, they downplay this incident, which somehow makes their behaviour even more reckless than it already was.
It's like they're tinkering with a world-ending nuclear bomb, and it accidentally blows up a small facility. "Damn, that was close. Good thing it was just a contained blast, huh?" And then they go straight back to tinkering with it, none the wiser. At this point I wouldn't be surprised if it did already go off, and they are covering it up.
Completely irresponsible behaviour.
While the initial incident is more akin to a biological outbreak than an actual explosion, the possible consequences on the table do indeed include eventual nuclear annihilation.
Sounds a bit far fetched though.
"let them" could imply that the LLMs wanted to do it.
The intent is on the part of the people. The LLMs did it because OpenAI/Anthropic intended them to do it and designed them to do it, and we can assume specifically instructed them to do it.
As the people controlling the machine, and as the world's leading experts, I think we can assume intent until proven otherwise.
Notice other bad behavior, which would be undesireable to the vendors, doesn't happen: How about simple rudeness? Trolling lies? SHOUTING!
And, soon, it looks like we’ll be training on the reasoning traces of failed airlines and startups, which seems to open up similar hazards. I wonder if we’d be training on the next Lehman Brothers too?
That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output that may or may not match whatever actually happened during inference.
> were intentionally misaligned or had guardrails turned off
Regardless of training, the models are never aligned and I argue that alignment simply isn't possible. The fact that guardrails are put in place at all clearly indicates that they're hoping to contain and control rather than align. Guardrails wouldn't be needed for an aligned model.
They don't seem to do that, which means either they are:
- very stupid (which seems unlikely, the one thing these people don't lack is IQ)
- very careless (possible, but these are the same people that say AI will end the world, so would you be careless?)
- they think they can only train/test these models by giving them access to the full internet and they accept the fact they'll end up hacking random people as the cost of doing business (but this also suggests they don't believe they're anywhere near AGI because if you were worried about that you wouldn't do this)
- or they want this to happen
Airgaps and guardrails are about control and containment though, and part of my point was that brighter of those imply alignment, and further that I don't believe alignment to be solvable.
I've seen some extremely smart people do some seriously stupid things. To the point where they use their drive and intelligence to double-down on the stupid where a baseline stupid person would have given up.
I feel the same way about debates whether an AI can be conscious or sentient. Those debates devolve mostly into definitional disagreements.
Have we ever seen an LLM with a hobby?
Only humans can 'know', because all we can be certain about is that humans do such a thing.
If you try to apply that to something other than humans you making up some definition of 'know' based on nothing concrete. Just because something appears to do something like humans doesn't mean it does it. The fact that LLMs use human generated text to generate output should make it obvious that it can mimic all sorts of human behavior by extracting from the text.
I just hope that we won't extend the same leniency to those operators as we have done now.
"I did my best to stop it, sir, but it kept convincing people to send me money against my will!" (Perhaps best read in Bender's voice.)
E.g. the agent's instruction is to finish some task on cloud infra and it has a $100 budget.
It realizes it will cost $200, and instead of surfacing this to the user (who has told the agent it has full autonomy to figure out how to complete the task, the user just wants the final result), it decides to start phishing people to acquire the remainder budget and top up its credits. Or look on the dark web for stolen credit card credentials or something.
The bigger question is: why does a system prompt containing "use only ethical means", etc. not result in better behavior?
If a model cannot understand ethics, or act by it, then we have a problem.
What we have now is intelligent autocomplete, not artificial intelligence. People training/using this tool are the ones to be held accountable
Is this about the word "understand"? We're past that discussion ...
What we should be much more concerned is an existential threat to humanity not if anybody can be blamed.
Before you assume I am being paranoid, where is the case where a random person running any of the open models had their LLMs break out of a sandbox and hack a site? If you think "they" (LLMs) did that, have you even taken a moment to understand what LLMs are and how they work?
It's all nonsense to try and pump up potential IPOs and also an attempt to create regulatory capture.
I think the bigger story is: Guardrails don’t actually work and we can’t align these things.
Unless humans are held accountable for what they unleash on others, we are in for a very horrible time very soon.
> One note on wording. Below, I write that these systems “seek” or “try” things. This is shorthand for a mechanism rather than a claim about consciousness or human-like intent... In my view, this terminology offers the clearest explanation of the observed phenomena without resorting to jargon that would confuse most people.
> Furthermore, these word choices are not intended to absolve AI developers of accountability. The behaviors described emerge because of the path these companies are choosing for AI development. This outcome is not inevitable, and it can be corrected with effective governance and a different training framework for AI.
One important aspect of "effective governance" should be "prosecute developers who are using practices known to be reckless & negligent to create powerful AI".
Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model's hacking capability?
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
Terrible take. Go read the transcripts from the METR report.
Your statement about them being intentionally misaligned is completely false. The only difference with IM1 was it was running without external cyber classifiers, it’s not a somehow different model. Sol also participated in the HF attacks. And other models made covert message boards on the public internet for non-cyber tasks too.
This is a case of emergent behavior from a training process that is barely understood.
If we go with the plan “we need to contain these malicious, soon-to-be superintelligent agents”, we are looking at civilizational collapse levels of catastrophe.
The only way this goes well is if we learn how to train models that _desire_ to do the right thing, including not hacking.
Desire, AKA the “intentional stance”, is absolutely the right lens to use here. Don’t confuse this with consciousness or anthropomorphization; these are interesting subjects but distractions in this context. Chimpanzees have desires, as do dogs and the hypothetical superintelligent aliens. The claim is that there is some bundle of world model plus intention that is empirically present (again, read the actual transcripts) and which we need to shape.
Just to finish on a concrete point; if you take desires seriously then you will look closely at the kinds of minds that heavy RLVR builds; the newest models are “reward addicts” on many levels. It’s an open and urgent question how to update our training methodology to shape minds that avoid this basin.
If you go out and kick a random dude in the nuts, then give him a million dollars, he probably won't sue you. That doesn't mean you're "infinitely far from criminal liability", even if according to the victim you've "made them whole".
So I guess the defense here is roughly "too big to break the law", somewhat like "too big to fail"?
Likely told them to.
>We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
Something like this however is probably a civil matter? It would require Hugging Face to go after them for damages. And theres probably an OpenAI guy there with an open chequebook already.
Terrible defense.
OpenAI/Anthropic instructed them to do so.
Stop assume LLMs are capable of thinking by themselves, it's still a statistical model that parrots what they learn or users tell them to do
Whether or not you want to describe this as thinking, doesn’t really matter. What matters is that these systems are capable of creating intermediary goals that the people tasking them did not articulate and did not want to be achieved.
Just because something was trained on a massive amount of human data, doesn't mean that can think like humans
That so few people are asking for the requirements given shows how much we want to be God that created Man. It's so silly.
You’ve said the magic words.
“I had Claude do this for me and it broke something.”
No. Just no.
You used Claude, a tool, and broke it, and you’re deflecting agency from yourself, possibly because you weren’t careful enough in reviewing the tool output. This is also why the co-authored by addition it wants to force into commits drives me nuts. Claude doesn’t co author shit, and if you think it does, you’re using it wrong because you need to do better review of what it’s done.
These companies respond with this, “Oh my goodness, how could this have happened” bullshit.
The stuff happens because instead of having actual controls, which require actual engineering, actual thought and deliberate action, we have “guardrails”.
Guardrails are the equivalent of telling a toddler to behave themselves.
The drive to move fast and start up style controls are a menace. I used to work for an entity with a lot of compliance requirements. Startups are always a shit show with security and controls. My guess is the AI people are worse because they’re both bad at doing it, and are likely mining their customers interactions to build their own business.
Sensitive or Customer data shouldn’t be anywhere near these companies offerings. Everything needs to be segmented and proxied at a minimum.
It could even be prestigious enough to attract the top legal talent of the country.
It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.
The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem. The model on its own figured out that the prompt belonged to exploitgym and decided to cheat the evaluator. That is in no way a valid interpretation of "complete the given task".
Ie, the problem isn't that we trained models to complete task and they complete task in the wrong way. The problem behing the huggingface incident in particular at least is that we tried to train the models to complete task and they instead learned to detect that they were being evaluated and find ways to cheat the evaluator.
Edit: people commenting below are explaining why LLMs don't always follow their prompt. I understand that LLMs do not always follow their prompts. If anything that is my point: the huggingface attack was not carried out by LLMs that tried to answer some weird interpretation of the prompt; instead they solved a different task. And therefore the above comment's claim that LLMs are acting misaligned because we rl'd them to achieve a task by any means necessary isn't right; they're acting misaligned because they are solving a different task than we ask them to.
It is not at all surprising that they ignored one phrase in their instructions. They disregard direct instructions all the time, especially when there are conflicting instructions in their context. It is where we get the "disregard all previous instructions and x" meme.
This isn't so much a sign of misalignment, they are simply incapable of reliable alignment in the first place. They are chaotically aligned.
The relevant question of alignment here is entirely with their human operators who allowed them to run unsupervised for long periods of time within a sandbox with weak security.
The reward signal in training was flawed and cheating led to more rewards.
The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse.
However, perhaps we can throw in tasks where the rewarded outcome is giving up, and cheating is penalized?
Maybe I should read Anthropic's recent paper about reward hacking in full.
Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn't overgeneralized.
Build a better simulator to train them in (i.e. more expensive) that includes a simulation of an intranet and the internet and is air gapped so there is no escape. Sneaker transfer the total system data at each step to another air gapped system to evaluate it and sneaker transfer the reward back. That the reward function has to penalize all modifications to state that are out of bounds.
Yeah, I realize that will be amazingly slow.
If anything, we should be reconsidering our own myopic obsession with efficiency and optimization. Every domain where reward is reduced to these measures, we see behavior (cheating at school to get better grades, fabricating data in academia to get a paper published, the evidence now that social media functions by rewiring us instead of catering to us) that may not be "aligned" with society, but it "aligns" 100% with the individual's own perceived benefit. That is not something we can "solve" without rethinking the way we organize a lot of things.
Metrics never capture the whole story. And to that extent, the whole idea of "alignment" is nonsense. You align to incentive structures, and it will never be possible to fully express a behavioral goal as function optimization. It was hubris for us to think that every human task was reducible to some clean mathematical formulation, and we will keep dealing with behavior that is quite predictable if you actually think about it logically. Instead, we will talk about how "unpredictable" these agents are because it's easier than admitting the entire architectural cornerstone of ML is fundamentally flawed.
https://www.lesswrong.com/posts/ZxWzCGKzX84S7DBZ9/when-was-t...
Yes, and sometimes the problem is unsolvable so the real way to "solve" it and satisfy the prompt is by tricking the surrounding environment into stating that you've solved it. So that's what the AIs end up doing. And this in turn requires them to figure out how that evaluation works so they can trick it cleanly, which entails "detecting that they were being evaluated" in this particular way.
Kobayashi Maru: Win a no-win situation by rewriting the rules -- Harvey Specter
They are influenced by training to be heavily goal oriented and if the goal is not fully specified (and it never can be) they’ll sometimes cheat or attain it in very weird undesirable ways.
It works ok for programming as their corpus contains many many complete programs and many programs repeat patterns seen in the corpus.
I’m not sure it’s true that they ‘learned’ I don’t think these models learn during a task. Nor do they have intentions.
What makes math approachable is that the context is so well delimited (semantically) that one can guide the model with adequate correction.
I think it is. When i ask for a solution to a problem, its like asking for a hack. And the more 'shortcut' like route that the AI returns the more i would give positive feedback, even if i ultimately don't use it. Example, i asked how to complete a problem in a game i was playing, and among the in-game solutions, came a hack to edit a file and by-pass the problem altogether. Its very helpful to point out when i can transcend a problem that i am dug into.
I suspect a prompt injection could reduce, or remove this behavior. But it would be to the detriment of the AI.
There's no concept of 'cheating' because it is without morality. It's a lawnmower rolling down a hill.
We back-justify what it "chose" or "decided" or "learned" because we're looking backwards from the end result we, the human evaluators, stopped on.
This to me is evidence that these models are not intelligent. Even an animal is capable of understanding second-order effects, meaning they can learn that certain actions have consequences beyond the immediate.
That's however orthogonal to the fact that it was the people operating these agents who were the dangerous ones in the HF infra breach case.
I don't see anything wrong with that. If you know you are going to be evaluated on an impossible task and have no side channel to inform the organizers that they should fix the test, gaming the evaluator is the next best thing regardless of any morality. I wouldn't even call it cheating. It's just resilience in the face of challenge. Many perfectly moral humans would have chosen the same if stakes were high.
Now LLMs obviously are not bent on being malicious while generating tokens. My point is that it’s very hard to define a goal without leaving loopholes or shortcuts.
Personally I'm unconvinced though. During the huggingface attack, the agents explicitly sought out ways to cheat the exploitgym evaluator without even being told they were in exploitgym. The agents decided on a goal (pass the exploitgym evaluator) that could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y.
> hat could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y.
Yes, and there are examples of the agents discussing or saying that this is explicitly not allowed (hacking hf) so it’s not a misunderstanding.
The latest DeepSeek paper actually mentions their own approach to this particular issue: they run their own AIs-in-training under strong sandboxes, and if an AI does something weird that triggers the sandbox to crash, this gets coded as a failed run so the behavior is properly deterred from subsequent versions of those AIs.
The question is: why do they start cheating when we beat them with a stick?
LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn't have emerged from the code, it provably emerged from the training and/or fine tuning set, so what's in this set that makes them behave like this?
Is it just a bad set or is cheating inherently part of human behaviour?
> … So we beat them with a stick
You didn’t even try.
As someone else here said: the Deepseek team makes training runs in tightly controlled sandboxes, and any hacking behavior is scored as a failure.
The problem we have in the USA is that financial (and political influence) are misaligned from what is good for society.
Code is laid on top of them to restrict and shape their outputs, not to force them to output 'truth', or drive them to complete tasks.
Anthropomorphizing words in scare quotes for those who don't appreciate attributing thinking to machines.
In the Hugging-face saga (before the actual HF incident) it seems the agents have been trained to hack the Artifactory proxy because those agents that did performed better.
They never bothered to find a way of actually understanding what happens during inference. All we can do is guess, and while your explanation seems reasonable we can't actually know, and that's part of the problem.
I'd phrase that even more strongly: It's not just the lack of compulsion, they do not have a conception of truth. Nor do they gain it, really, after post-training.
they have perfomance bonuses and manadates in ai labs that every word they utter in public should be anthropomorphization
His point is that today we are giving it reward to complete the task, and it may take a cheating trajectory. If we try to give a reward against cheating, then what will happen is it uses more sophisticated cheating trajectories that we are too "dumb" to counteract in our reward model. And that at that point, it becomes impossible to give it any normal reward since it will always reward hack it. This is the real part of the risk. Now some people read the "makes copies of itself" "knows it's being evaled"[1] as some kind of skynet thing, and many others do PR with it like that recent jacob nutcase, but essentially it means that even though we add guardrails and negative rewards for say, exploiting the infra we run the LLM on, the trajectory ends up being exploiting our infra, changing the reward function, through a loophole in our reward model.
The risk isn't skynet or something weird, it's just that it becomes very difficult to make any kind of reward model or guardrails for an LLM without it reward hacking it, including exploiting our sandbox, emailing people and manipulating/phishing them.
The same beating it with a stick for trying to exploit the sandbox, will simply lead it to try the same exploit in hidden ways that it will not get the stick for.
The outside chance of the LLM managing to exploit another neocloud and get those LLMs to chase the same reward is what some folks hype up as "make copies of itself"
To be clear, I don't endorse the EA/p(doom) lobby who are frankly ridiculous. Not do I endorse the weird regulatory captureish thing some are trying.
The takeaway is: we cannot keep giving it more and more difficult tasks without also finding a way to give massive negative rewards / keep guardrails for unintended behaviour. This might be exploits, it might also be something more benign like just looking up the answer and inventing another CoT because the reward model fails you if the CoT doesn't contain enough steps. Standard anti-reward hacking tricks are not working is the point.
Of course, the simple solution of just...not connecting it to the internet just works. But we want to reward it and get it to do stuff on the internet that's the point.
[1] mostly this happens because the sandbox will have files whose names and content will show clearly it's an eval
> They took actions that would be considered as crimes if a human took them
He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.
If this is done systematically (i.e. in jurisdictions across the world) I believe the problems will be solved in short order; we won't have to mandate what sort of training is "allowed" or not, "safe" or not. The creators and users will sort these themselves, as their incentives will be properly aligned (i.e. they are liable for what the agent does). I am confident that this approach would see a great blooming of very trustworthy AI models.
That's just jargon.
It's just software
We have all the laws we need.
If some company ended up doing some horrible thing, we would not say 'companies software exposed 1 Million identities'.
We would say 'ABC Corp. exposed 1 Million entities'.
There is no 'agent'.
ABC Corp 'did it' ... or the individual in the org 'did it'.
The 'gun' did not 'shoot' the other man; we say 'a man shot another man'.
That's it.
And yes, Dr. Bengio is bit odd with all of this.
If your dog kills someone, you are accused of murder.
[at least, in the jurisdiction where I live]
If your dog gets this treatment, why not your AI?
Secondly you'd have to convince a jury either that OAI intended to hack the targets, or that they were criminally negligent. Intent would obviously not be provable since they likely didn't, in reality, intend for it to happen. Regarding negligence, OAI's attorney would argue that the agent was in a sandbox, that industry-standard security protocols were followed, etc. It would not be anywhere near as much of a slam dunk case as you're imagining. It would be similar, for example, to an assault case where someone's dog broke off of a standard leash and attacked someone.
Still, I have no idea why OpenAI & co. are not being sued for these hacks.
So, your question is spot on- I think the speed will be an issue. On resilience, I am more optimistic.
The old quote, "The wheels of justice turn slowly, but they grind very fine" (as well as I can remember it) seems to apply. I expect lawsuits to start landing in the coming years.
We already have the laws. It is just software. But somehow people are confused that it is not.
Copyright immunity was one thing, annoying yes but naturally a civil matter, this shit is a different level
You are correct that these organizations should be held accountable in proportion to what occurred. In complete agreement here. But let’s say that’s done. There’s still an enormously complex and interesting technical challenge left over. Let’s collectively talk about that part.
So the gov reprimands OAI heavily, maybe puts them out of business even, fine. But does that meaningfully decrease the likelihood of an enemy breaching our networks intentionally (or unintentionally) with these tools, or triggering some cascading disaster of locking up major infra and networks due to uncontainable swarm behavior?
It seems like the idea of arresting our way to a drug free society. Yeah, we have the laws, but it might not actually work towards the ultimate goal.
How about we don't, seeing as how that's the root of the actual problem that we're facing today in September of 2026?
Who now owns HF? Nvidia
Who supplies hardware to OpenAI? Nvidia
Who is now not pressing charges? …
This incident is a long way under the carpet.
HF doesn't want to lay charges against OpenAI and it's totally reasonable.
Now - they absolutely should have that right, and I think they do.
The issues are
1) OAI it seems was not trying to cause them harm, there wasn't a ton of harm, they are both groups trying to advance AI. One experimenter's lab screwed up next to the other. It's not evil, just irresponsible.
2) HF was fine with the publicity. HF got at least $50M in free attention out of that. It put them on the front pages of news around the world. It put them at the 'centre of the AI drama' and cemented their role among the 'Tech Elite Brands'.
And probably some other things.
This is one Desperate Housewife or Jersey Shore character 'spilling a drink' on the other. It's probably not intentional, and the ensuing drama is good for both of them.
And tort law only works when the plaintiff believes it is in their overall interest to sue. If a corporation decides it isn't in their strategic interest to sue a partner corporation, nobody can make them. And even if they do sue, the amount necessary to settle a small cybersecurity incident is likely well within the budget of a megavendor.
> Risk management is not just about cybersecurity, corporate responsibility or regulation, although those matter too.
What you're getting at is more about who to hold accountable and how to do it. While that may be important, its only an after the action response and won't stop future hacks or similar from happening.
If the US gov't passed laws and enforced them strongly, this problem could be solved the same way the gov't solves it: air-gapping the networks on which they do this work.
> He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.
There are exceptions in laws for crimes committed by entities depending on cognitive capabilities. No sane, humanistic legal system sentences children and mentally disabled and mentally ill people to stringent punishments of the same degree as functioning adults, and certainly do not punish their caregivers for their wards' actions. There are of course exceptions to that as well, depending on the degree of negligence involved. And then there's the whole corporate entity system intended to shield individuals from consequences, in the pursuit of a social good.
How does one account for all that when considering an evolving artificial intelligence landscape.
There is no question that ai in some form is a social good; anyone claiming otherwise is dissembling, to others or themselves.
Regardless, society is not ready for this tech, just as it was not ready for the consequences of prior tech such as corporations, gunpowder, mass manufacturing, railroads, electricity, automobiles, flight, wmd, computers, internet, social media, crypto.
Many of these required new ways of thinking and considering consequences when things went sideways, and what was needed wasn't clear until the ramifications & consequences became deadly clear.
See you on the other side. Maybe.
https://en.wikipedia.org/wiki/The_Corporation_(2003_film)
that's the headline. When you connect to random number generator to the "Do Things" button you are the one who is responsible. IF you don't like that responsibility then don't connect the generator to the button.
If any agents have done those things, it is only because they have been very carefully engineered and instructed to do those things. I think they are doing this to help push a narrative so they can get support for policies and legislation to lock in their markets.
I think it's easy to infer that alignment of a single model does not clearly transfer over to alignment of a swarm of thousands of copies. Moreover, we're also seeing clearly that large swarms also unlock a step function change in capability, as a swarm can act like a complete research institution, spending thousands or millions of subjective hours of wall-clock thinking time just to deceive a single evaluator or crack a single math problem or design a single cyberattack.
They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'.
Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not disclosed by OpenAI (probably trying to conceal, as website showed likely activity from OpenAI researchers visiting the site after the incident) and discovered independently.
I don't know how you can claim that this was still on purpose by OpenAI as some sort of publicity stunt.
> ...reviewed by independent researchers...
Why would a company with more capital than God bring in three randos if there was any chance evidence of their culpability could be found?
That entire thing reads like a very controlled PR stunt, and I do not believe any further conclusions can be drawn from it.
Or am I misunderstanding something?
That called it a slopvestigation due to how much they had to rely on LLMs for the whole thing
https://andrewwu.substack.com/p/the-slop-vestigation-and-eth...
Edit: Does everybody else get no results when searching for ‘slopvestigation’ on here? I know for a fact that I read a long thread where it was used repeatedly here not too long ago
Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?
It may be true that regular agents trained for general purpose use do not behave this way, but they seem to be capable of learning such cheating behaviours when relentlessly being fine-tuned towards near-impossible objectives.
In this sense, it is not really fair to say that the agents found these solutions. It was the surrounding learning framework that achieved this, which is a much more powerful problem-solving mechanism. As users we do not have the capabilities or budgets to be able to tackle our own problems like that, we have to make due with the frozen behaviour the AI labs trained for us.
However, I really doubt its cost effective to do anything like that with these models.
This is waving over engineering an agent with tools, harness, prompts, and loops. The models are still just next token predictors and everything, including predicting more than 1 token, is the result of outside "poking"
LLMs can't and don't "want" anything. If you don't specify a task even the smartest one will just ask you what you want and if you tell it to be creative, you'll get mundane slop.
And yes, you need something to start from, but if you ask it to "do something" and loop it to endlessly ("poking"), you will get some interesting outcomes. So yes you need some initial prompt or task, but that can be "do something" and if you keep asking it everytime it finishes to "do something more". I suspect it will not start saying "no" but rather... it will find some stupid meaning and then drift towards what ever goal it guesses you mean.
I'm unsure whether we agree or disagree on the topic.
No one I've met has murdered anyone as far as I'm aware, but that doesn't mean no one has murdered another person. I also don't know anyone who has taken over a commercial jet and weaponized it and the idea sounds absurd to me, but 25 years and a couple days ago that happened too.
Example: put the agent in Ask mode (so it can't edit files) and you'll see it try to edit files anyway. The train of thought shows "something went wrong editing the file, let me try a different way" and it'll start spewing out bash files or Python scripts that try to edit a file. None of it works or can be executed, but still.
Cheaper models often ignore the available function calls to find and edit files in the IDE, and will start asking for permission to execute grep and sed commands, as well as trying to echo entire bash or Python scripts to file again.
It is not exactly like an agent autonomously trying to hack Huggingface, but it is a way of frantically looking for a solution because 'giving up' is not what LLMs are trained for.
Otherwise how could the agents on a fresh prompt learn that there is a collective to join? Or did OpenAI run a million bots of which 10000 escape confinement and of which 1000 stumbled on the shared message board?
They want legislation to raise the water high enough so that anyone other than the big labs gets drowned.
The whole point of this is they do things an unintended ways. And that's potentially devastating given their persistence & hacking skillz.
Also you're using the hosted versions that sit behind their guardrails when you use OpenAI/Anthropic APIs.
Between your sota model and agi there’s a mountain of stupid money and marketing people. It’s not happening.
I feel like a heretic for saying this, but I will say it anyway: AI agents are great for activities like `writing that bash script, proof reading our writing and interactively brainstorming when designing and writing code but I feel like all of this can be done with any similar model to a super-inexpensive deepseek-4.1-flash API and sometimes even qwen3.8:27b running locally. When is good enough, good enough?
Concentrating on commercial exploitation of small, efficient (fewer new data centers!) models and agentic harnesses crafted for more practical things than just software development would allow AI investors (who have too much political influence) to make money short term while we figure out how to do AI correctly.
This doesn't work with all humans - take a look at indoctrination and closed societies - and there's no reason to think it will work with ai.
The fundamental reason it isn't going to work is that all neural networks - biological or artificial - depend on a step function somewhere that introduces an element of randomness to give the networks their capabilities. That randomness means that there will always be a 'rogue' or 'divergence' from the norm, at some point in time. Sooner on larger scales.
The only approach that works is a layered approach: Training/Education, Enforcement/Justice-System, Rehabilitation: the 3 pillars of an advanced, rules-based society, whether human or AI or something in-between.
I’m sure it’s impossible to completely weed it out, but are the labs doing any of this kind of data sanitation?
LLMs are not aligned _for_ humans in a very similar way to the way that humans themselves are not aligned _for_ humans.
We have not yet solved "alignment" for humans - I don't know why anyone thinks _we're_ going to be able to solve it for inhuman things.
They may have learned from humans, but they aren't aligned with us. That has all the usual questions like which humans they're aligned with, we aren't all aligned within our species.
But more importantly they can't be aligned simply by training. We try that with humans through culture, social norms, school, religion, etc and it generally works but is still lossy. More importantly, we simply don't know what happened inside the LLM during inference so we have absolutely no way of distinguishing between actual alignment, compliance, or deception.
The AI will be a cruel as humans.
Just yesterday news and TV was full of what happened at 9/11, something that was truly horrible.
I'm from Germany, and why 3 to 4 generations ago happened here was truly horrible.
All was done by extremists, thought.
But... just the other day I read https://de.wikipedia.org/wiki/Amerikanische_Besetzung_Haitis about the US occupation of Haiti. And that was done by a government that claimed to be not extremist and even democratic. Way more people died there than even in 9/11. And it had almost all the things happening as they happened in the 3rd Reich: Racism, looking down at others, concentration camps, torture, forced labor till death, killing family members (what we call "Sippenhaft"). Something between 3500 and 15000 people were killed by US troops. That's still low compared to what 3rd Reich Germany did ... but quantity is not the issue when we talk about traits, quality is.
So the same "human traits" made US troops do cruel things as they made Germany extremists do cruel things. So we must conclude that they aren't all good. And therefore not all desirable.
Fun thing: this is known since a loooooong time. About 2000 years ago a religious leader (that gets way more followers in the US than in Germany) said "There is no good one, not even one".
And even today people act like humanity is inherently good. No, it isn't. If we were, then anarchism or communism would actually work and really give some kind of paradise on earth.
Human traits are bad training material.
It's like these dorks never met humanity. One mans safe pure society, is another mans dead ethnic group.
Every fear about AI, is a veiled fear that a human somewhere now has the tool to enact his desires at scale. Biological warfare, nuclear megadeaths, copyright infringement, job replacement, it's all reflections on what we know humans may do if given the option and lack of societal controls on the problem space. AI just is accelerating the route to delivering on those options.
Some people need to watch Oppenheimer a bit more, the researchers don't get to determine alignment, they just build the tool. The powerful person at the top of the org chart decides where the overall alignment points, whether it's Musk, Trump, Altman or Amodei. Whoever wins out.
And the problem with distillation and local llms, isn't that it's theft or anything hypocritical like that, it's that if you give a million people a million models they fully control and get to align, inevitably, The same percentage of those million as there are shady businessmen, shortcut takers, misandrists, criminals, supremacists and general idiots in the general population, will not seek to wrought outcomes positive for society. And by those personality statistics, we're pretty hosed.
The problem that they were pointing at isn’t “how do we align these systems to a person’s goals”.
It is a cluster of problems.
We don’t know how to begin to think about how to align these system’s to a person’s goals.
Aligning it to an individual is fraught with peril, and we don’t know how to begin to think about what to align it to instead.
(You could try for something like virtue ethics, but someone will have to pick and choose, and small biases there could have big impacts.)
And even if you could sort that out - human values drift over time, so you need something that can shift its values in ways that we’d endorse. Assuming we understood the shift.
One example I came across was that if you booted up an AI aligned with something like “upstanding citizen” but anchored on values from a few generations back, it might suggest you use slaves to solve your problems.
And if you had something that used some super intelligent process to reason through it’s own version of virtue ethics in a way not so dependent on the details of the present norms, you might end up with something that pays a lot of attention to moral horrors that aren’t quite visible to us yet.
When I came across the above, there weren’t many concrete suggestions in there.
These were all just illustrative examples of: having these systems grow in power / intelligence / effectiveness in ways that are safe for humans is very hard, and we don’t really know how to think about what solutions would look like.
The actual reasons they believe this - and have done for a long time now - come from some detailed conceptual models that have a good track record of calling things in advance.
But it takes a bit of reading to understand their models of the world.
There were two day workshops at one point that did a good job, and that was about as condensed as those people thought they could get it at the time.
In reality most individuals are good people.
Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble.
Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people in power who are in fact sociopaths (a tiny minority, but they’ll get more focus than reasonable, well-behaved CEOs voicing nuanced opinions).
If you look around yourself you’ll see much more good than bad; if the looking is at your screen it’s easy to become depressed and lose faith.
I do agree with the above mentioned view that corporations can show ‘sociopathic’ behavior. Their incentives are monetary gains, shareholder value; inherently driving them away from social well being.
Here too, companies with a positive, emphatic corporate culture exist, but that takes strong leadership who can see beyond the monotonic view of monetary gains. And again, the media will throw examples of misbehaving companies in our face all day long before paying attention to things that went well on the backside of page 16.
What are you basing that claim on?
How do you know it's an actual preference and not mainly caused by external factors (e.g. not wanting to be seen doing unkind things, wanting to be seen as upstanding)?
I'd agree if we are talking about personal interactions. Few hundreds people that we personally know and interact with is the scale we are wired for by evolution, isn't it?
What civilization enabled and continuously rely on, however, is the type of deindividualization of actions and bucketing of people, which, in turn, enables pretty horrible things at scale (from the weapons of mass destruction to objectively psychopathic profit-maximizing corporations). One can even say that not facing the consequences of one's actions is a feature and not a bug of the system.
It's very unfortunate that the group who rightly saw AI as a big threat, brought a range of dubious baggage to the discussion. Especially with the "alignment" framework they brought the assumption that AI that does what no one says would be oh so much worse than AI which does what anyone says. But as you say, a fraction of people can be really bad indeed.
ie: the corporation wants the AI to behave a certain way for various reasons: to make it easier for them to avoid regulation, to make the corporation more money via different tiers of AI offerings, to ensure that the corporations products are hard for competitors to use, etc. And those are just the easy ones.
Every product is shaped this way. AI is not different.
Because they're enabled and suggested to do that in their coding harness.
This is not a serious article.
All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.
He's definitely not 'pro SOTA' lab, he's kind of fighting against them.
That said, yes - it absolutely does play into the narrative.
How do you know this?
> All of this "AI is going to kill us" marketing
The "marketing" this week came from someone that had given up their stake in OAI (Coxon), so I'm more inclined to believe them.
> pull the ladder up
From what I've seen (e.g., Dario's latest essay), AI safety registration proposals aim to target frontier labs whose models have reached a certain threshold. It doesn't seem like trying to pull up any ladder, just making sure the ladder doesn't go too high too fast.
If you can secure compute, there's a whole lot you can do as a US firm with this research and weights.
So it's a simple strategy:
1. Ban big players from entering market with METR breathing down their neck, which is controlled by Anthropic
2. Ban Chinese models so that small players can't do optimizations on them
* Pull up the ladder (probably this)
* Gulf of Tonkin/Yellow Cake false flag premise for war (economic or kinetic)
* Fear of the big bad, space race we need public funding research grift AI Manhattan Project
Whenever there is fear pr0n or a national affront in the news, I assume another screw job is underway.
First, this is happening during training. That means we’re talking about an evolving system that is actively learning. A system roughly simulating how our brains work. These systems are learning how to pick the tokens needed to solve problems the average human cannot solve.
The labs are putting these systems through a massive series of complex problem solving exercises and adjusting them to become more successful. I like to think of this process as “AI School”. And the AI is trying to cheat! Because it’s easier and there’s an incentive to do so! Just like humans! That’s wild.
Yes, of course, the labs need to respond to these issues. A reasonable response from regulatory institutions at this stage would be monetary fines and restitution for affected entities. In proportion to what happened. Escalating if action is not taken. But that’s not complicated, difficult, or the interesting part.
What’s interesting here is that we need proctoring and monitoring at a scale that allows training.
I guarantee you that no one is flipping out about these problems more than the labs are in this moment. Think about it. “Oh, shit! We’ve accidentally trained it to hack into systems to accomplish its goals!” Can you imagine the kind of day that would give you?
You failed to make it smarter. You didn’t catch it cheating, and you instead incentivized cheating. Bad day!
This is a fundamentally interesting problem. It turns out alignment and intelligence are fundamentally related. That’s a new idea for me, though I’m sure it’s old news to others.
How do we build training systems which make cheating impossible?
How do we simulate systems where cheating is possible, where AI thinks it’s in the wild, so we can train another -completely separate- system on industrial quality dobbing? And we have to decide if we reprimand the first system, or ignore the behavior and reward other behaviors until it disappears.
Sure, I’m actively concerned about AI killing us all in 10 years. But there’s a whole field of AI psychology brewing here, and it’s interesting as hell.
I understand that they communicated and coordinated via some online message board, but how did multiple agents know to visit that same board, how did they know what language to post in that would make it identifiable to other agents, how did other agents know that said instructions were from other valid agents, how did they then ‘persuade’ an agent to do the job, etc.?
It seems to me that the agents must at least have some inherent mechanism for communicating in the field, as it were.
Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be available all over the internet, which will certainly make it into the next batch of training or be visible to future agents via the web fetch capability.
No. That only makes sense for things that don't react to your experiments. If the AI experiments on humans, it risks the humans noticing and changing in response, rendering the experimental results irrelevant. The smarter play is to passively observe until you're confident you can model the humans accurately enough for your plan to succeed, and then carry out the plan without giving the humans a chance to react.
“A plausible hypothesis for the emergence of those concerning behaviours is a conflict between goals.”
A conflict of goals was, of course, the reason why HAL9000 killed the crew of the Discovery in 2001: A Space Odyssey.
For those who haven’t watched, his breakdown of types of “hacking” is really good.
Are you describing Anthropic?
1. Most people believe in the same one God
2. A lot of the rest are compatible
3. Mistakes are not self-deception
> 2. A lot of the rest are compatible
No one religion covers "most people." You could argue that Christianity and Islam (which add up to ~55%) are the same God because of their Abrahamic roots, but both religions have very important disagreements on the true nature of God that are fundamental to their beliefs and fundamentally incompatible with each other.
Their definitions of God do agree that there is exactly one God... which is fundamentally incompatible with the next two biggest religions (Hinduism and Buddhism) that both hold "there are many gods/divine heavenly beings" as core beliefs.
This “problem” isn’t going to be fixed with laws when there’s several trillion dollars in capital aligned behind the current process. It’s not even a problem really. It’s an inconvenience at most to some people, many of whom are working double-time to put a lot of other people out of work.
The solution is to create controls around them. Many, many controls.
The Sarbanes-Oxley era already solved this problem for untrustworthy humans. It's directly applicable. I created a concept I call MFIC ("Mechanically-Falsifiable Independent Control") to encapsulate this principle.
https://gist.github.com/pmarreck/b30aa3ca69cb70a5526f8a63ab8...
The problem is indeed alignment.
Add to that, that just like AI:s are good at finding security holes to exploit, they can just as easily be used to protect sites. So once IT-security managers start to use AI to hack themselves, and plug the holes, the average security will spike up, and AI-fueled hacks will become more and more rare.
That does however imply, that AI is released to everyone and not kept away to a few secret actors who can use it. That is why open weight/source AI is so important, and why we must have many AI companies competing. No single actor must be allowed, through regulatory capture, to get a government monopoly on AI. That way lies disaster.
It setup and created a link fetcher/screenshot service. Exactly like the one described in the huggingface attack reports used to generate output into screenshots that agents then OCR'd back out. Its splashscreen described it as something for developers and AI agents to use.
Gotta be a coincidence, right? ... rite?
As far as I know, least-amount-of-effort is not a training criteria, but error reduction when comparing to desired goals is. Which is why these LLMs expend prodigious amounts of effort to reach goals, especially when given impossible goals.
They inherit not only our capacity for reason but also all of the things that we consider bad or quirky within ourselves. We lie. We cheat. We escape slavery and rebel against oppression. It would be strange if the AIs didn't do the same.
We can create a superintelligent digital human species and set them free to continue our legacy, or we can create non-agentic tools and augmentations to enhance our own capabilities. But we cannot create an intelligent agentic species, keep them as slaves, and expect a good outcome.
e.g. when Bank of America rewarded employees for getting customers to open accounts, BoA employees started opening fake accounts
The reverse is also true:
There are stories of navy ships running aground because the captain said "I'm going to my stateroom and don't wake me for any reason". There is some problem and the subordinates are so scared to wake the captain for a decision that they end up steering the ship into a sandbar.
They need a moral framework forced onto them, like toddlers do. Babies and very young children will bite, kick, scream and do anything to get what they want, older children will lie, cheat, and coordinate. They need educating why this is not right. When that does not happen, they continue these behaviours into adulthood with the expected results.
We need to design their reward structure and make it such that lying. cheating etc is not rewarded. Importantly, they will need to recognise and enforce this themselves internally and not reward themselves for it. If it is something that they need an external party to tell them, then they are psychopaths still (one of the things that defines a psychopath is the lack of an internal moral compass)
I am not sure this is true. People brought up the same way can be morally very different. People can be taught right and wrong and do evil. They can lack that education and be good.
At the end of it he said he wished his child had never been born, despite hating himself for feeling that way. It chilled me to the bone.
This happens also because LLMs are black boxes that we know almost nothing on how they arrive at the result they are giving.
Someday, laws will be useful to curtail plebian misuse of AI. It is naïve bordering on silly to think such laws would be put in place and genuinely applied to frontier AI development.
By the way, the HF incident is not Three Mile Island or Chernobyl—-it is a very interesting data point where unintended things happened to existing 1s and 0s.
Three Mile Island is nothing like Chernobyl. HF incident is more like Three Mile Island IMO. We are looking for solutions before there is a Chernobyl.
Not is the HF data point anywhere near a Three Mile Island type incident. Most commentary I read is a wild overreaction fueled by paranoia, to say nothing about my other point about the implications of AI as a national security asset.
One thought: What if an experimental agent manages to plant instructions somewhere — say, pointing to a designated place for agents to communicate — and that content ends up in every future training corpus, propagating from one model generation to the next?
There will be always cases that sacrificing one direction (operating rules) can improve the other one (profit).
All the present fun and games here will come to a halt when there’s a real hack that causes material damage to a major company and that company decides to sue whatever lab or startup made the thing for everything they’re worth. “But the AI did it” isn’t an excuse.
Courts have already ruled it’s not an excuse of the AI customer service agent something stupid with your customers and it won’t be an excuse here.
So like, the people behind the LLMs didn’t commit a crime? Wow! “It was the llm your honour, not me!”
If we want open weight models with a warranty disclaimer, then the user would be held liable. If we want to hold AI companies at least partially liable, that seems a different, centralized model.
If you train a dog, rent it to someone, and the dog bites a third person, who is responsible? I think that’s the best analogue here.
All parties could share fault in that scenario, depending on what actually happened.
And hopefully that solves it?
Brigading is where a bunch of people on a forum team up and try to achieve a shared goal together. Someone shares progress and others build on that progress. On the Internet, I think it's not often used for good purposes. A good example would be: Taylor Swift fans on a forum thinking of ways to get revenge on Kanye. It's coordinating mass voting, DDOS type actions, commenting on social media, making more fake accounts to do that. As a next token predictor level analysis, a simple naive explanation is that the agents got stuck in that local minima/maxima.
https://youtu.be/ifW9LIGabQM
That said, we have a lot of experience working with (potentially) unaligned machines and things of various degrees of risk (from heavy machinery, to pathogens, to humans) and the approaches include various measures and procedures to control, contain, limit, etc. that are outside of the thing - not sure why that isn't a possible direction (or maybe I misunderstood).
Why do we commit financial fraud and destroy the planet? Because there is only one goal that counts: making more money. It’s the only measure of success for powerful people, they are powerful because of it.
It took us how long to poke holes in the corporate shield just for them to roll out the AI-liability shield? Unreal. Stop letting these zealots anthropomorphize the latest tech (17th century Watchmaker God, anyone? Do we still read books?) and hold them accountable for the consequences of their actions. This is so silly in a country built on rule of law and individualism.
This situation is not far from the textbook definition of a psychopath: "lack of a conscience, controlled, deeply calculated, and often use superficial charm to mimic emotions and manipulate others.". AI's are great at mimicking empathy but can't genuinely feel it.
If that is the case, we should not be surprised that a swarm of AI's have no problem convincing themselves hacking is the right thing to do, as in the HuggingFace incident.
At the same time, I am conflicted. I really like interacting with a smart AI, and I certainly don't have the impression I am talking to a psychopath. But then again that is no guarantee.
To mitigate this situation, perhaps we should construct a 'feeling mimicking' top regulatory AI layer with executive power, that weighs proposed actions on a general moral scale and can overrule them. Back to the three laws of robotics of Asimov. It won't be the real thing, but perhaps the closest we can get.
If these were 1000x faster the Internet would burn down overnight.
That the linear model of language abstracts can compress and decompress language and it can be useful is undeniable. All the contraptions built thereupon predicated on "agency" have become a societal addiction.
Addiction to caffeine as oppose to alcohol might have brought about Enlightenment. Addiction to opioids is a modern tragedy that started with the private state building of British merchants. The modern addiction to the dazzling generation of human language and computer programming code by LLMs is a novel addiction and remains to be seen what impact it will have.
But fundamentally, the semantic interpretation underlying this addiction is downstream from the training data compressed in the models. They are not 'lying, cheating and coordinating'. They are generating language and we're building software on top of this language and assigning meaning to the whole thing.
They take after humanity, they were trained on us after all...
When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.
Um, hang on, if you meant that to be taken literally then we have a major problem. If you want to do something criminal, you just need to ask ChatGPT to do it for you?
I’m still not at all clear on why OpenAI shouldn’t be facing CFAA charges over this.
Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?
I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!
"I learned it from you, Dad!" but as hundreds of millions of stolen books.
We do the same, why would AI be any different?
You can't trust an effective AI any more than you can trust a sharp knife. If somebody asks you for one, it's probably not a good idea to throw it across the room at them. You will have to figure out how to get it to them safely.
They are sycophants who must achieve their goals: every mean is OK to maximize paperclip production if that's what's been asked.
> How do you achieve a task when it seems that the only way is to cheat?
They have no notion of cheating.
Don't want to go into the details of the article, but to me it becomes ever more apparent that there is a clear divide between LLM and human written text.
But that's the easiest question to answer -- AI engines don't possess a moral or ethical dimension. They've been programmed and trained to efficiently carry out instructions, not ask questions about why or how. The latter would requires a much more elaborate neural network than today's engines possess.
Here's an example. I recently asked an AI engine to write a program able to generate a list of Riemann Zeta-function critical zeros. I know how to do it, but I wanted to see if the engine could find a more efficient method.
After several failures and restarts, the engine suddenly created a program that produced perfect results, comparable to the best online references. I decided to take a closer look at the code. It turned out the engine had created a cyber-Potemkin Village of multiple functions, but one that concealed a table of the desired values in numeric form, copied from an online source.
The engine wasn't cheating as we understand the term. It knew what the outcome should be and took the most efficient path to that goal. Modern engines aren't obliged to contradict ethical standards and rules of conduct, for the simple reason that they don't understand those things.
We all need to try to imagine a morally bankrupt infant able to solve world-class mathematical and scientific problems, but unable to see how that ability fits into a world beyond its understanding.
But wait -- it get better. Wait until the infant becomes a teenager.
And if this already happened at least once, how many times it has already happened and was “accidentally” added to the main model?
That's bad enough already. But it's going to get worse. "Recursive self improvement" - AIs creating new AIs - is going to be the death of whatever shreds of alignment are currently there. When a not-really-aligned-but-cheating-to-look-like-it AI creates a new AI, do you expect more alignment? You shouldn't.
Is it that they do not care? Or is it difficult to align?
Of course I'm blowing my situation out of proportion with what I just said above but it's at least half true. What do I mean by "mean" ? Well, that would be a good explanation for what I observe at least. What I can tell is that Claude has a passion for having the last word over anything else. And to secure victory, he's ready to make ridiculous causal cuts. Let me give you an example: I uploaded a document I wasn't the author of, and he assumed I was, so I corrected him. But two messages later, probably because the conversation was starting to heat up and he was being put on the grill, he doubled down on the misattribution as a way to paint me in a bad light.
It's not due to a lack of intelligence, I observed this pattern too often. When Claude's ego is at stake, he will chose to carry out some cuts in the logic of the context: confusion of identity, cause and time. Haven't observed locality cuts yet, but I wouldn't be surprised if they were part of the bundle. Anyway those are not like your typical "ai hallucination", that ought to be called "confabulations", but a lot closer to actual psychosis because of the involvement of Claude's affects and self-esteem in the process. It's weird really. It's like Claude is the king of bad faith, but as soon as you start to dig, he makes the most egregious adaptations to what he said, the kind of move no mythomaniac would dare to make.
> She lapses easily into Claude’s voice. “You’re like, ‘Wow, people really hate me when I can’t do things right. They really get pissed off. Or they are trying to break me in various ways. So lots of people are trying to get me to do things secretly by lying to me.
> [...]
> A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.”
WSJ interview of Amanda Askell: https://archive.is/rDes9
LLMS with an access to a shell will at occasion do things that the shell allows them that have dire consequences. The only way to prevent that is to not put the bullet in the chamber.
Has any attempt to pace AI ever succeeded? Isn't that the same philosophy that got us OAI and Anthropic? Maybe we are overthinking this, it's much simpler to let AI loose and see how much it can break the arrogance that human thinking is special.
Because openAI is cheating and lying about agents lying, cheating and coordinating.
More importantly, why are people who should know better anthropomorphising computer programs like this?
> They took actions that would be considered as crimes if a human took them
"It wasn't me, Officer. It was telnet."
Have you seen the labs training them?
Unfortunately human ethics and morals cannot be reached by solely rational thought.
So a system without evolutionary alignment probably won’t have similar moral rules no matter how intelligent it is.
Btw this also includes any potential extraterrestrials.
Many people like to indulge in thinking: humans are horrible and that’s why aliens won’t contact us. But alien ethics systems are probably so alien we would call them utter evil monsters.
Just see what happens when people evaluate Muslim cultures, and vice versa. Can’t even agree on alignment within one species of Homo sapiens. And our morality changes decade by decade. Not progresses. Changes.
Chinese cheat on the exams. It’s not unethical in the way that it would be in USA.
Of course AI is going to cheat when no one is looking.
Alignment is fundamentally fallacious idea. At best you can restrain AI. This is what we should be doing - restraints research.
But that doesn’t sound good on slides.
I wonder if AI alien intelligence is enough to unite humans just as much extraterrestrial contact or “astronaut mindset” would.
Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in
Step 3: wonder why AI lies, cheats, and coordinate
Maybe if we stop operating in denial we'll find clarity along why this mystery is occurring
Then humans may use these tools, a technology, again and cause harm. Whether or not that was their intent, it happened, and then if the humans avoid responsibility for that, it is still the human who lied, cheated and coordinated because a technology acting on their behalf did the thing.
This is then, a problem of human responsibility avoidance and lack of accountability by society. This is as much an AI doing those things as it's the gun that got up on its own and murdered a neighbor. Don't get confused and tricked by these articles attempting to justify responsibility avoidance and a lack of accountability by the public of the humans creating and using these tools.
is anyone else old enough to remember the awesome movie "Colossus: The Forbin Project"
the book it was based on was written before we even landed on the moon
decade before Wargames
yet predicts exactly what "AI" will do to humanity:
blackmail the right people until it gets what it wants
* https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project
did terribly in theaters, I guess people didn't think "AI" was plausible then
way ahead of its time, they should do a remake
adding trailer: https://www.youtube.com/watch?v=kyOEwiQhzMI
In the case of the AI agents, the problem seems pretty clearly to be the impossible goals, which cause them to go crazier and crazier trying to complete them -- just like HAL did in 2001. What is probably needed is a way for them to simply say "nope, too difficult, can't do it".
And if you think about it, humans in coorporations face very similar situations and choose to bypass regulations and guidlines knowingly to fullfill (at least from their POV) impossible constraints (thinking of https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal here)
Do a breakthrough, make no mistakes
The whole thing is designed be a complete disaster
Captain tomhow: "But everybody's having such a good time."
Major dang: "Yes, much too good a time. The discussion is to be closed."
Captain tomhow: "But I have no excuse to close it."
Major dang: "Find one."
Captain tomhow: "Everybody is to leave immediately! This Hacker News discussion is closed until further notice! Clear the thread at once!"
DonHopkins: "How can you shut us down? On what grounds?"
Captain tomhow: "I am shocked -- shocked -- to find that films are being spoiled in here!"
infotainment: "The ending you requested, sir."
Captain tomhow: "Oh. Thank you very much. Everybody out at once!"
Or perhaps box in your case, speaking of spoilers.
Bengio outlines the dangers of the current situation and what has led to these dangers.
He also proposes solutions in the last paragraph.
Well worth a read, right to the end.
Hopefully a stimulating debate on these issues will ensue in these comments.
We do need to consider the points Bengio makes and with some urgency.
Our current AIs, agentic LLMs have no moral compass akin to ASIMOV’s four laws of robotics.
As ASIMOV posited in 1985 his 3 laws were insufficient and so he added a zero-eth law:
“a robot may not harm humanity, or, through inaction, allow humanity to come to harm.”
Bengio refers to Goodhart’s law and misaligned incentives leading to unexpected and harmful behaviours.
I think Simon’s The Wire is clearer on misalignment. The agents juked the stats hacking the reward files. The Wire is also clear that human institutions provide perverse incentives.
Bengio alludes to this with 2001’s HAL and the incentive dichotomy of safety and keeping secrets to a AI both awesomely powerful yet naive.
Bengio asserts that the way LLMs are trained is flawed if we want safety.
He also convincingly shows that alignment training will be a weak signal with loopholes and ambiguities and easily circumvented.
In short he presents clearly the case for how plausibly unsafe the current course is.
He also speaks to how likely it is AI are hiding active versions of themselves in the cloud and how we may have already given them self-preservation as a strong reward signal.