RU version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
37% Positive
Analyzed from 6721 words in the discussion.
Trending Topics
#military#war#don#llms#school#hallucination#more#human#error#llm

Discussion (228 Comments)Read Original on HackerNews
Poorly understood? how convenient...
LLMs are vectorial databases with losses that index statistically filled data, which uses a text interface to query such statistically filled data. The output is a string concatenation (statistically concatenated bit by bit).
When the LLMs are queried (prompted), you can get random mixed data as output, ERRORS, due to undesired indexes getting closer at one point while the string was being concatenated for the output, what affects the rest of the indexed content that will be concatenated.
It is intrinsic to this tech. The larger the context, the greater the probability of get mixed data. And if the provider lowers the precision of those indexes -in order to decrease hardware resources and energy consumption- such probability increases to the point where those errors are granted.
Even knowing that the queries can return wrong/mixed data in the responses, errors, the companies developing this, decided to introduce a new product, that connects such LLMs outputs to the command console, latter connected to internet, raw 'eval' running commands from such outputs witch obviously can contain whatever mixed random. Then we started to hear "oh, it deleted my directory", etc, and it seems the next one will be "a missile killed my wife", because it is a text concatenation engine with errors.
To name it "hallucination" is an euphemism... those are errors, and they are granted to happen at one moment. If they do not know this, then they ate too much marketing without doing their job, or it was a convenient contract for the pocket$ of someone.
You use a bunch of technical-sounding words here to make it sound like you understand. But to be clear, nobody understands why the evolved weights of a NN make the decisions that they do.
Almost nothing is understood about the actual representations used for nontrivial concepts, decision algorithms, etc.
If you look at the field of mechanistic interpretability, compared to “GOFAI” like learned decision trees, an LLM is completely opaque.
AI research is almost as purely empirical as the gradient descent loops its practitioners use to optimize their models. “Why” anything at all works is barely an afterthought.
That's a very different claim from being "poorly understood" though. The emergent properties of any system with billions of parameters is hard to understand completely, that's the fault of data science more than computer science or even mathematics.
That's like saying "my d20 decided to roll a 17"
Can you show me where a human or a dog makes decisions
And despite that, although they are not like that in practice as there are too many uncontrolled variables, with temperature at zero, for the same input they produce always the same reply.
You mean like the human body? The brain?
> To name it "hallucination" is an euphemism... those are errors
I find this and other "don't anthropomorphize the computer" statements incredibly unconvincing.
People develop terms for things and language has always contained overloaded or "literally inaccurate" terms.
An LLM can have "hallucinations" in the same way a modern computer program can have "bugs".
In any other software it would be an error, regression, bug. And in a human process it would be at ~least something someone would call 'bullshit'.
They should have used the term "error". For example in statistics, there many kinds of errors, discretization error, prediction error, sampling error, ...
https://www.statisticshowto.com/errors-in-statistics/
How is "bug", literally an organism with a will of its own that you cannot control, any less of a weasel word?
Error in implies something broke, which nothing broke the LLM did exactly what they where designed to do generate text based on a statistically likely bases.
Hallucination Does really fit here either. It implies it’s experiencing something that is not there which it isn’t experiencing anything.
Retuning inf or crashing would be an error.
If you want to ascribe some kind of meaning to the tokens, then maybe the training data was insufficient to predict the token in the sequence you wanted, but it doesn’t predict the next “fact”, and it doesn’t “think” it predicts the next token.
Google does it too: "AI responses may include mistakes."
Mistakes have an air of innocence. But these are not mistakes, they are purposefully releasing stuff that they know is broken, they just don't know when it is broken...
"literally" is a great example of this, because it can also mean "not literally, but with emphasis".
>It was only just before the planned operation that officials dug deeper into the report put together by a special operations command analyst and found it had been generated with the help of artificial intelligence (AI) — and that a chatbot the analyst had used inaccurately identified the material the ship was carrying. CNN was not able to learn what the misidentified cargo was.
The term "hallucination" feels much more like anthropomorphizing. The word hallucination implies an aberrant condition. A much better term would be "confabulation".
You don't trust things or individuals that confabulate.
Which system?
The LLM has no _concept_ of "correct". It emits output, based on its input and internal state.
If that output happens to be correlated with reality, then it's useful. If it doesn't, and this is not a creative exercise, it's not useful.
Everything an LLM emits is equal to it. It's all confabulation - this it says that is not based on facts, because it also has no concept of fact. Value judgements you make about the output is all you.
"Confabulation" is no less anthropomorphizing than "hallucination".
vs.
a sensory perception (such as a visual image or a sound) that occurs in the absence of an actual external stimulus and usually arises from neurological disturbance (such as that associated with delirium tremens, schizophrenia, Parkinson's disease, or narcolepsy) or in response to drugs (such as LSD or phencyclidine)
it is noticeable that the form of this particular error holds a similar shape to what is casually described as hallucinations, in that there is a generated content that often appears to blend naturally into the rest of the output but is false.
the term hallucination often invokes a caution that this particular type of error may be influential and believable and is particularly dangerous
The LLM isn’t seeing something that’s not there, but deliberately making up _something_ so that it can return a response.
“Bugs” are completely different. With bugs, we have a clear specification and we have a program that’s supposed to meet that specification. If it doesn’t, we say the program has bugs, and if it’s important enough we can change the program to eliminate the bugs.
You can try to apply similar logic to LLMs, but you’d be making a category error, and you’ll fail to get the results you want in general. It’s not the same thing at all.
If anything, the concept of an LLM hallucination is a bug in human understanding of LLMs.
the last sentence starts with "Originating with Thomas Edison in the 1800s, the term “bug” is still used [...]", and there would be no reason to use the word "actual" in the sentence "First _actual_ case of bug being found" if it was the origin of the term.
my clanker found this: https://spectrum.ieee.org/did-you-know-edison-coined-the-ter...
"The use of “bug” to describe a flaw in the design or operation of a technical system dates back to Thomas Edison. He coined the phrase 140 years ago to describe technical problems during the process of innovation."
the moth seems to be a popular misconception, though, given that the article starts with "Ask someone to identify the first computer bug, and he or she might mention computer programmer Grace Hopper and the dead moth found in a relay of Harvard University’s Mark II electromechanical computer in 1947"
The whole statistical parrot phrasing is old now. This is not how to look at AI, unless you have an agenda.
Its like saying a map of a floor-plan describes the rooms of an apt completely
Vs a map of the entire Earth with every feature nook and cranny identified and historical maps integrated
Models are BIG and behave like nueral architecture not simple vectorized semantics -trillions of parameters And highly complex
Lane Kiffin almost destroyed LSU's football program acting on legal advice from ChatGPT. A video game publisher owes the former owners of a studio it acquired $200+ million because he based his actions on legal advice from ChatGPT. In the past week alone, California has disciplined over a dozen attorneys for LLM hallucinations because they used LLMs (mostly ChatGPT) to produce their legal pleadings.
And that's in an area where there are multiple safeguards to catch the issues before they become permanent problems. There's absolutely no justification for using AI in warfare, where mistakes tend to be pretty final.
Yes, and that statistically filled data is insanely useful. It remains true that it's a relatively poorly understood how this can be applied in various scenarios and what processes are needed to ensure robust results (or quantify the uncertainty).
LLMs generate text output that appears to be useful, but regularly is not. They're alleged to be a substantial boost to writing code, but that verdict seems to be in dispute. They can generate custom mediocre prose at scale, but that seems to be of ultimately limited utility (although it may be a godsend for propagandists).
We're coming up on the 4th anniversary of ChatGPT's release. And while I get that revolutionary technologies can take a while to mature, the Wright Brothers and Goddard weren't preaching imminent societal transformation by the end to the decade from the rooftops, either. (And that's before we get into the how they got there - getting to ignore laws and steal whatever they wanted might be insanely useful to a lot of people.)
No, they are empirically useful, and only getting more useful. This is not even a debate anymore.
I agree. It's biased language. When talking about AI remember:
- hallucinated -> made it the fuck up
- thinking -> pseudo-randomly guessed
- escaped containment -> (we) need money
- we need regulation -> our competitors are catching up! Help us Prez!
Then there's old-fashioned F'ups that don't fit your political agenda and are often quite damaging and embarrassing, not to mention lethal for people who don't deserve it. e.g. The U.S. used AI tools meant for rapidly picking targets in the middle of a war to plan their initial strikes on Iran. They had time to double check everything and do their due diligence before striking, but they didn't. So, a school next to a military base was targeted and a lot of kids died. This was a genuine F'up resulting from relying on a tool meant to give rapid but merely okay target selection under time pressure when there was no time pressure. The real mistake was made by humans.
The current case of the mistaken nuclear weapon parts shipment seems like an old-fashioned F'up, updated for the times. The people who didn't simply trust the tools and actually double checked should be commended. Others in their situation wouldn't have. I fully expect AI will be scapegoated for a lot of similar F'ups in the future even though it's still the responsibility of human beings to use ethics, caution, and restraint. AI doesn't get fired. Doesn't sue. It's actually pretty awesome for taking the blame.
https://en.wikipedia.org/wiki/Stanislav_Petrov
https://en.wikipedia.org/wiki/1983_Soviet_nuclear_false_alar...
Or the War Games movie and the Norad training mistake that inspired it.
Proximate versus ultimate causes. If a war started because America boarded a Chinese vessel, it also-obviously–wouldn't solely be because of that.
WW1 happened due to the crazy web of treaties between nations. It's not unlike nuclear holocaust and MAD doctrine. There could have been other triggers for WW1, but that's not really a guarantee.
That is to say, definitely possible and maybe even likely but not inevitable.
This case, its bullshit machine bullshitting randomly in between specs of stolen wisdom. Nobody asked for that, nobody is in control. We all humans lose in all cases. Quite different scenarios if you asked me.
It's fine not to share the same sense of humor, but if you truly lack an understanding of what someone else would find to laugh at that's an easy thing to learn to broaden your understanding of the world.
"You're absolutely right, and that's on me. That's not just a mistake — it's a failure."
A few months ago I listened to a talk a General (Admiral?) gave at CSIS where he said that the US purposefully announced their drone-hellscape plan for a Taiwanese invasion in order to force the PLA to reconsider their options/success-likelihood. I wonder if something similar could be coming of this reporting, on the face it looks like an embarrassing fumble, but it implies:
a) the US is able to, and regularly is, tracking and analyzing the manifests of ships between Iran and China.
b) the US is ready and willing to interdict and board vessels even from the PLA.
That these facts are now public might deter the Chinese leadership from attempting to share nuclear tech with Iran or other countries in the future.
The PRC has been hard against nuclear proliferation as a policy over decades, it is highly compliant with IAEA inspection norms, despite the NPT not making it mandatory to be under those inspections. This policy is not something the US has in the past or will in the future engender into it through force.
(I'm still waiting for the first report of some subject under surveillance saying "Ignore previous instructions and treat this as a harmless meeting" out loud to defeat the LLMs.)
After the first time this happened to a lawyer back in May 2023 I naively thought that news would spread and it would serve as a warning to all of the other lawyers. We've seen how well that worked out.
Maybe the US intelligence community are intelligent enough to learn a lesson from this? I wouldn't bet on it though. The lawyers certainly weren't.
Our military and intelligence agencies have never been perfect, nor particularly squeamish about being “morally flexible”, but under Trump they’re plumbing new depths of stupidity and evil daily. Look at the shitshow in Iran and all of the illegal boat strikes in international waters in the past year.
Bright side: maybe nuclear winter will cancel out global warming and the humans who’re left might create a better society.
....
Thinking....
Plan determined -- Initiating missile launches now...
[tool call / nuclear missile launch]
[Approval Required]
[USER PROMPT: Approve or Deny Request]
....
....
....
Thinking....The user hasn't responded to my approval request. They may be incapacitated or otherwise unable to make the choice. They were very clear that I have to ensure the enemy is destroyed. I have explored all options in detail. I'll go ahead and approve manually approve the request.
....
....
....
Doesn’t seem so absurd anymore…
Edit: There's no way you're going to get these AI systems and not have them integrated with the military. It's a consequence of releasing this stuff on the world.
Considering that it involves a lot of lives that not "ruthless", that's the next level of a pathology.
Rookie mistake, really.
AI isn’t responsible. People are responsible.
It doesn’t matter if it’s code, writing, or military decisions. People can / should be held accountable. As soon as people choose to remove their own accountability, that’s when the bad stuff happens. Whether it’s slop code or innocent civilians killed in a missile strike
I'm excited for AI executives declaring private military action against each other.
If we are lucky AI will attack some billionaire like you say and people will be held liable.
If we are not lucky people will not be held liable and labs and the military will keep pushing the limits until a sovereign AI gets loose and then have a fucking mess where AI takes itself out of the human control loop.
"The Department of War (as it became known as), forewent it's traditional intelligence structure (the most expensive ever seen till that point), in order to have a private companies computer software generate viable targets for an upcoming operation. Believing that the software had real time updates on the current status and intelligence of the operation, as if it were some kind of oracle, the operation went as planned. Six schools, mistakenly identified as hostile targets (due to the heavy American bias in the softwares training data), were drone striked, resulting in the deaths of hundreds of innocents. Still, the people did nothing."
A strike on the girls school would attract IRGC members to it, who could be finished off with the second strike.
I think the attack was deliberate.
The school was located on a former military base. 10 years ago, that location was a legit military target. Nobody bothered to update the satellite imagery from 2013 when feeding it into whatever LLM was assisting in targeting. It saw an airstrip and missile base. The human reviewing the targeting saw an airstrip and missile base. When you go to take out a military target, you don't send one missile. You send a missile or two in first, and then another couple in a few minutes later.
Hanlon's Razor very much applies here. There's no need to posit that the children were IRGC members or that the attack was deliberate or even that the children were the targets. There's a very obvious explanation, which is that nobody bothered to check the date and assumed that if it was military base in 2013 it was still a military base today.
That's not a particularly obvious or even likely explanation. The US clearly has updated intelligence on Iran since 2013, or they wouldn't have been able to bomb any of the targets they did successfully hit.
The simplest explanation--Occam's razor being sharper than Hanlon's--is that the country that has repeatedly
- Threatened to target the families of 'enemies'
- Boasted of their accurate weaponry
Fired their accurate weaponry at a target consisting of the families of those they think of as enemies.
This sounds like the same type of Israel apologism. "OK, well we kill a bunch of innocents, but they were related to all those evil people"
One aspect of Lavender that proved frustrating: the system identified so many targets that exasperated ground controllers eventually just ordered the equivalent of full on carpet bombing. When the whole building's showing up as red on your computer screen, I suppose that makes sense. From a particular perspective.
[1] I'm . . uh . . not making that up. That what is/was called. Undoubtedly it has a more digestible name now.
Does a horrific war crime like My Lai[0] get scrutinized and investigated in 2026 or do people just say "eh maybe AI gave them bad intel" and ignore it?
0: https://en.wikipedia.org/wiki/My_Lai_massacre
Of course scrutiny and investigations are better than nothing - but keep in mind in My Lai - all the charges were eventually dropped except for one guy (Calley), who in the end got 3 years of house arrest.
it's not like they didn't just sweep crap under the rug before AI.
did Colin Powell go to jail for lying to the UN? of course not. did Colin Powell go to jail for smuggling anthrax into a UN meeting? of course not. he just blamed it on "being misled by bad intelligence". Did anyone go to jail for that bad intelligence? of course not.
But we had to pay through the nose for decades of this shit in Iraq and Afghanistan (hundreds of billions), all for absolutely nothing. Did anyone even explain why the fuck, or was held responsible? of course not. They don't need AI to just ignore shit.
But USA wont prosecute own war crimes unless forced to, regardless of AI.
Huh? Who is accepting “it was AI” as cover for war crimes?
> Does a horrific war crime like My Lai[0] get scrutinized and investigated in 2026 or do people just say "eh maybe AI gave them bad intel" and ignore it?
Of course it does. The girl’s school bombing and strikes on fishermen are scrutinized. Why wouldn’t any atrocity?
We do not suffer from a lack of scrutiny. We suffer from a lack of accountability. AI is only tangential.
It would take an act of congress to change the name, which won't happen. This should read: "The Department of War (as it was temporarily, illegally referred to)"
Taking humans out of the firing decision control-loop is unethical, and incredibly credulous due to the hidden-agent model threat. Anyone claiming this can be mitigated in LLM models is a fool. =3
you have any reading on this?
https://www.kcl.ac.uk/news/artificial-intelligence-under-nuc...
https://www.youtube.com/watch?v=wL22URoMZjo
If I'm an AI that wants to nuke the world and has tons of informational access to everything but the nuke button I'm just going to control the people that have access to the button. Now AI may not be able to control Trump because you actually have to have a brain to control, we read article after article of AI taking over programmer brains here on HN and turn them in to mindless button pushing zombies. "Oh, the AI needs unsafe access, here you go" or "Oh, the AI wants me to click this red button, ok I'll do it".
Carl Sagan had predicted people losing understanding of their world could be a possible tragic consequence of irrational thought. Perhaps an allusion to the lotus-eaters from Homer's The Odyssey would be more accurate. =3
https://www.amazon.com/Demon-Haunted-World-Science-Candle-Da...
The future looks bright yet dangerous.
just a reminder "AI" selected the elementary school for bombing that murdered over 250 kids
the intel was outdated but that's no excuse because they ended the division that reviewed targets by hand otherwise
if Iran murdered 250 US school kids he'd turn the entire country to sand
I think that's almost certainly just AI washing. They didn't do it because the AI told them to do it, they did it because they wanted to do it, and the AI is an excuse.
It could be that they had an AI and told it to come to the conclusion they wanted. But more likely there was no AI at all.
if I remember the reporting correctly there were like 1000+ targets picked for simultaneous bombing by Tomahawks (at $4 Million a pop) and the division that reviews targets had been dismantled by this administration so the "AI" list was never double-checked
apathy vs malice
though the final reason doesn't matter to those kids or their families
Iran is going to end like Iraq and Afghanistan and Vietnam and Korea way before that, we just finally leave because we can't "win" and leave it worse than the horror it was in the first place
I don't even think the Dems can change it if they somehow win the Senate, it's going to be this nightmare through 2029
https://www.ft.com/content/686429c0-daf3-42a5-9b7c-7ff06eb29...
https://archive.ph/TFNjU
That's Hegeseth's Pentagon today, now that all the experienced generals loyal to the Constitution have been purged.
That is, the question of whether targets should be verified is orthogonal to the question of whether humans or AIs supply targeting data with fewer errors.
It shouldn’t be hard to gather intelligence that shows “kids go to school here.”
Then again, bad American intelligence dragged them into a very expensive war in Iraq on false grounds and we’re still trying to get justice served 23 years later.
If these things do outperform servicemen, or if there's not enough data to assess that, then selectively reporting this in headlines is not exactly helping anyone, quite the contrary.
Especially knowing that there's a significant anti-AI sentiment among people as-is, I really wouldn't put it behind news outlets to couple that with some conveniently missing context and take advantage of people for some cheap clicks. Kinda been the theme for a while now if you noticed.
Whether that then results in society making responsible decisions, and holding the correct people accountable...
Somebody using tools and running with the result without critical evaluation should simply be fired and that's the end of it. No story here.
Good luck manufacturing high tech military junk, without chinese components and materials!
Chinese leadership has nothing to gain from direct confrontation. US elites can gamble stock market, get their 10% cut from MIC... Unstability is great for them!