FR version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
51% Positive
Analyzed from 6616 words in the discussion.
Trending Topics
#things#llm#don#llms#more#using#google#something#enough#still

Discussion (143 Comments)Read Original on HackerNews
A fascinating dichotomy has become apparent between those who trust LLM output and those who don’t and don’t understand why you would.
Surely if the machine you go to for answers regularly makes things up you would just stop using it? Perhaps people have to be burned by something really bad personally before they realise the limitations? LLMs are very convincing and persuasive.
Before we would find an intriguing post on the internet from years ago, and you have to verify it with additional research--it's easy to skip that additional research.
With a LLM when you're skeptical you can interogate it. One thing we know for sure is LLMs are quick to admit mistakes were made when interrogated, comically so. A LLM might not always recognize its own mistake, but at least it is available for easy interogation, unlike the forum posts of old.
Manual research from reputable sources remains an option.
In my experience every LLM out there is utterly useless and quickly defaults into "here are other concerts that took place around that time near that location". Google Search (ignoring the AI overview) is even more useless, as it refuses to show literally any webpage that's older than say 5 years. YouTube search is genuinely better than Google at surfacing old and grainy fan-made videos uploaded in like 2010, but also defaults into synonyms nonsense pretty quickly.
But, the search functionality of exactly one forum and three local news websites that I know have an archive that dates back long enough beats every single one of those abovementioned every single time. Three people are talking about their experience at a concert on a random 15+ year old forum thread? It happened. The tiny list of 5 or so (Google-hosted!) Blogspot blogs I have bookmarked? They usually have a photo of the ticket that Google Images refuses to show me.
Not only are search engines completely dead as a category, but LLMs are a shit replacement for them. "We" (okay, Google specifically) has truly committed a crime comparable to burning the Library of Alexandria. Everything older than a decade that wasn't properly documented on Wikipedia is just gone, never to be seen again.
The funny part of this is that Google search is intentionally bad at returning YouTube videos, presumably because some anti-trust action scared them into artificially ranking videos from local news sites, Facebook, and other ad-walled content ahead of YouTube videos. Seriously, go watch a YouTube video, then try googling its title with “video” appended to it, and see if the “Videos” tab of google search ranks it as the first result.
It's a great tool, but verify the important things (or do them yourself)
Signal to noise has taken a dramatic hit.
AI shouldn’t have any reason to lie. But its lies aren’t intentional. It’s just actually making things up and “hallucinating” when it pretends that an option or setting exists, or confidently claims something entirely untrue, and makes up a source to go with it. For something Google is willing to shove into the top of every search result it’s crazy the percentage of time the answer is blatantly incorrect.
The goal of LLM's, as they are marketed now, is to drive engagement and stickyness of products. A wrong answer is brushed off with a "Hey, you're right, let's try that again" - a response purposely designed to maximise the friendliness of the system and minimise the sting of a wrong answer. The fact that an LLM will not respond to the same question in the same way twice (i.e. the 'temperature' ) is because increased accuracy will not drive engagement and, therefore, increased accuracy cannot be allowed to get in the way of engagement.
Psychics are still in business. Although to be fair they probably don't have as much revenue. I think people really enjoy being told how smart and insightful they are, and how much they've really cut to the crux of the issue. This isn't the whole thing, but I think it counts for a lot.
Even today, going to a therapist and having 1:1 sessions is a rational activity, and even covered by insurance. What do we hope to derive from therapy sessions but some personal insight and improvement?
You know, I go to church, and from my perspective, sometimes the most difficult discernment for an individual is between "The Holy Spirit's message to us in general" or "the general messaging to everyone around us" vs. "what I receive and my personal interpretation of things".
When preachers and oracles and leaders are speaking in generalities and trying to get big followings and trying to appeal to the widest audiences, that's when it's most difficult for us to determine what God is really saying to us, in our own hearts; that special instruction for our own lives. We can't actually get that from an oracle, psychic, or any 3rd party. It really needs to come from our own well-formed conscience.
And it's the same with a chatbot or LLM conversation. We can pose questions, make prompts, and get spammed with tokens and walls of text. No matter how personalized, it's still up to us to interpret that, and extract nuggets of news that we can use. It always has been.
I'm reminded of that person who killed themselves due to their discussion with ChatGPT and their parent wrote their obituary using ChatGPT. I don't think it is enough.
We live in hope that people stop doing stupid things and are constantly disappointed.
The last sentence reflects a lot of my feelings on the first question. LLMs have a sort of weaponized take on the ELIZA Effect. The better their memory the better they are at playing to human social desires to be listened to in an active conversation. At some point it stops mattering if the answers are right when the answers feel right, but really, like ELIZA back in the day, so much of what makes it seem special is just reflecting your own writing back at you in a convincing and persuasive way.
Steve Yegge likened LLMs to slot machines. The human brain is very vulnerable to random reward systems. If you get an hallucination, just pull the lever once more.
People still respond to ads and political speeches.
I feel like the most pragmatic perspective is "trust but verify."
This is why they're so effective at coding: you can run the code yourself (or the test suite) to verify that it actually does what it's supposed to.
And maybe these people finding hallucinated results on Rachel's site are doing verification too.
This perspective I really don't get.
What has any of the LLM companies done to earn my trust? I lean more towards "verify because I don't trust".
Not necessarily, namely because P != NP. Verifying the correctness of a solution is faster than solving it. Thus a system that outputs 99% incorrect solutions and 1% correct solutions can still be incredibly useful.
not saying anything new. easy enough to frame it like any other assistant and check references
I'll temper that slightly by saying it's mostly out of morbid curiosity because the things that the Dreaming Piracy Robot comes up with are frequently wildly incorrect code, but it's interesting to think about how it might have got there.
And then I think, well, maybe Special Needs Wintermute has a point. Maybe there's a different way to think about it that I've missed.
And then I just change it back to what I wanted in the first place.
It's painful to watch my older colleagues use their agents, and they're not even that much older. Like they were intentionally trying to sabotage themselves sometimes.
They're getting better, but the time it takes for them to pick things up is just significantly longer, not the least because they're kind of just throttling themselves in addition.
Good thing that there's not much to pick up on at least.
When I ask for their input, it's always for a situation where I'm capable of judging if their input is useful or not.
In all situations I use them, it doesn't matter if they're correct at all. I'm asking for ideas, alternatives, links for blogs or articles. I talk things out with them...
I don't think we should ever "trust" LLMs. This seems like the wrong usecase for them.
One smart engineer seems to have entirely offloaded all thinking and conversations to one, with just occasional editing. It's utterly bizarre to hold any conversation with him. It's one kind of rude thing if he was doing that to respond to me reaching out to him if he felt I'm not worth his time. It's a other when he's the one actively reaching out and asking my help with something.
When Linus posted that AIs and vibecoding were here to stay and declared resistance to it as harmful, I stopped to consider whether I was wrong, but it has made me realize that in retrospect Linus Torvalds and Linux itself aren't actually the holy grail of computing. I didn't feel that way with Richard Dawkins, its not like falling for an AI psychosis retroactively made me question The Selfish Gene, but now I'm looking at linux and the theory that it's a clusterfuck is gaining so much traction, especially after copy.fail and ensuing rustification, I see so much clearly now. It was never about linux, UNIX sure, POSIX, yeah, GNU fucking aye, kernel? Ok whatever, drivers and scheduler with a gajillion lines of code I guess.
1- Maintainers are allowed to commit LLM generated output.
2- Criticism of LLM generated code is not welcome/will be ignored.
Now, whether that constitutes being pro-Vibecoding or pro-agentic engineering, whether it's delusion, whether it will have problems, that's subjective. But I feel that whatever way you look at it, it's a topic that polarizes engineers, and Torvalds is taking one side and not the other. It doesn't seem to me that it's a very neutral stance, although it may be more neutral than projects like Bun or OpenCode of course, if it feels neutral, it's cause the overton window is shifting.
Edit: Guys, why are we downvoting this? Does no one use like ChatGPT or Claude and understand how it works? Do you all think its regularly hallucinating links still? Is everyone on HN using like free signed out accounts or something? What year is it?
The other day though I was seeing how well it could pull details of its own conversations with me. It often does this pretty well for broad strokes of things - it remembers, largely, what cameras I have and use when I ask photography questions. It's never made things up here, but it does forget details, such as whether I've bought something or am just considering it. However, when I asked it for a specific interaction I thought I remembered, it gladly went along with my false memory and provided an affirmative answer. It was the first time I'd been caught in a serious hallucination with a frontier model (Sol High on the web chat interface) in a long time.
When did that stop? May 7th, 2026?
I feel like a Google Maps-style system would discover this automatically by noting via phone location data that there is heavy traffic in Chinatown.
(I do get the author’s point, but I think that factually this example would not be a problem)
This kind of thing happens regularly.
I think it takes someone (at Google) manually marking those roads as unavailable before it will stop trying. I’ve seen it happen with other things too, like if a highway is closed because of a bad accident.
Accidents on highways are a different story because there are rarely any alternatives, or if there are, they're so much slower that it still makes sense to suffer through the 30 min jam.
Google is able to estimate crowd sizes in a business or on a public transit vehicle. This seems to be based on the number of Android devices reporting their location in a cluster. Maps could obviously make inferences if there were large crowds of people, not moving in vehicles.
I don't see any reason for this to be true. "Actual understanding" (which I take to mean something like a predictive world model) and desire for self-determination coincide in humans because of our evolutionary history, because our reward function involves reproducing in a competitive environment. Artificial systems usually have a very different reward function. IMO the burden is on the claimants to show why these two imminently separable concepts are likely to co-occur again under wildly different pressures.
So I’ll be handwavey here and say that if “actual understanding” is to an LLM what an LLM is to a bash script, so there’s a mechanism there that doesn’t just follow hardcoded paths, it takes new data and processes it in novel but human like ways to come to a new conclusion if needed, then the author is right, in my opinion.
We don’t know what the reward function is for an AI but AI is trained on so much human work that it probably starts off with the same biases in its understanding and reactions. It feels like there’s something very basic in becoming more independent the more you understand of the world. Animals go through this as they grow too, not just humans (listening to parental authority until they eventually don’t anymore).
Since there is no established consensus on this one the burden is on either party to prove their own side. Just because the author said something first doesn’t mean they need to write a full proof while you get to say “nuh-uh” and that’s enough.
They really shouldn't.
Does nobody remember why fizzbuzz was a thing? People who talked a big game while having no actual competence or understanding of the subject matter?
Nah. Separate things are separate until proven otherwise.
So much bullshit appears in the form "X is Y" where X and Y are related but distinguishable phenomena. If you really try to pretend that every such phrase deserves serious consideration just because it feels plausible to someone, you drown in nonsense immediately. (Of course what you actually do is only grant that consideration to ideas that feel plausible to you personally, but it's obvious why that's not a good principle, right?) So we have to make all of them justify their existence.
Anyway:
> So I’ll be handwavey here and say that if “actual understanding” is to an LLM what an LLM is to a bash script
No. None of this. Pretty sure this is just an incoherent analogy.
> so there’s a mechanism there that doesn’t just follow hardcoded paths, it takes new data and processes it in novel but human like ways...
You're putting your conclusion in the premise, right there in the open.
> ...then the author is right, in my opinion.
This still doesn't follow from your handwaved premises as far as I can tell.
Indeed. C.f philosophical zombies and the Chinese room problem.
The idea that some “actual understanding” (or more precisely, the outward appearance of it) needs a mechanism that includes free will is a bold claim.
This is an extremely reductive way to look at human existence. So much so that I read this with Richard Dawkins voice in my head.
Our existence is far richer then just the capability to reproduce. We (as well as other animals and even plants) do far more things then multiply, and in fact we often do things which are detrimental towards the prospect of reproduction.
I think it is actually a mistake (philosophically speaking) to try to find a simple reward function for the human existence. I see no reason for such a thing to even exist (let alone be simple enough to summarize in a single sentence).
I think however OP is correct that in assuming such a link exists, then an intelligent agent will trend towards preserving its own agency while also preserving its own existence.
That said, I am a skeptic when it comes to intelligence. The term is fraught and IMHO not a helpful term in science nor philosophy. We are better off simply defining intelligence in anthropocentric terms and claim no non-human system can become intelligent by definition. I would even go further, given how much racist pseudo-science the quest for explaining intelligence has produced, we are better off abandoning the term entirely.
I see the same thing with LLMs in software development. If you say "find a bug in this code" it will regularly confabulate bugs. If you ask it for a test-case, run the output through some deterministic thing that tries the test-cases, and tells the LLM it's wrong, the output of that system will mostly be legitimate bugs[1].
For now, transformer-based generative AIs seem at a minimum like a very useful tool for dealing with "squishy" problems when you have some way to validate their output. Many of the 404's to the blog are probably people validating the output of generative AI, which is the opposite of the inference made in TFA.
1: It will also occasionally hack your test-runner; I suppose that's also finding bugs, just not in the software you wanted to find bugs for.
That's the horror buried underneath all the tech, policy, and gloss. The real, animal brain, desire that drives most of this is: I'd like a slave I don't have to feel bad about.
We want things to do our work for us so we can do other things. That's not bad, and it's certainly not the same as literally enslaving another human.
It's a really complex machine, but it's still a machine.
I think people (many/most) don't want slaves.
Our imaginations just outstrip our abilities and we desire them to match.
1. False-negatives, where relevant posts that do exist are not being shown (imagined or otherwise) to the user.
2. Posts which exist but don't fit the words the chaos-parrot uses to describe them.
3. "Relevance" being determined by unpredictable factors that aren't stable, predictable, or desirable.
In other words, it's just more whack-a-mole lipstick-on-a-pig third-animal-idiom-here.
[0] A bad idea on its own, since it creates a security vulnerability for data-exfiltration or indirect malicious attacks.
Always comes to mind when I see Elon and friends getting excited about AI robots. Slavery was more about economics than the role-playing.
The whole gig/services economy is just building this up piece-by-piece: you can now pick the set of household needs you want taken care of for varying levels of money; and practically everyone participates in one form or another. This is exactly a disaggregated 21st century version of servants: paying for convenience. Of course with many issues in implementation, but I don’t see the ethical/moral issue with wanting this kind of thing?
https://www.nber.org/papers/w31758
https://www.noahpinion.blog/p/nations-dont-get-rich-by-plund...
Race might be a decent analogue to the difference between human actors and sentient AI, but I suspect work animals (plow horses, etc) would be a better analogy.
The premise seems to be that models aren't smart enough to understand this, and if they were, they'd be sentient and want autonomy.
For an article that's about making things up, and being too trusting, this seems bad. Maybe the author knows a lot about LLMs, but it doesn't seem like it.
Pasting the verbatim quote from the article into a free ChatGPT session: https://chatgpt.com/s/t_6a5e6f5e24f08191b6a482aad63cae63
Going to an incognito window and using a less leading question: https://chatgpt.com/s/t_6a5e6ee4a3508191bc1b352b41911b53.
If I go generic and just ask if there's anywhere I shouldn't drive, it doesn't get to Lunar New Year until I ask about "events" on the third question: https://chatgpt.com/s/t_6a5e6fb239208191b18cebcf7642c8b0. It's sort of a win for the article, if you think that people who run driverless car companies are all dumb, and won't create a prompt to tell their LLM to "consider events that might disrupt traffic."
Well the answer is actually that it's always a bad idea to drive straight through the middle of Chinatown at any time of the year, because the streets are narrow and full of tourists.
As an aside, what’s the problem with the extra traffic? Perhaps she has a lot of traffic but nginx can return a 404 with a tiny amount of CPU.
0: https://wiki.roshangeorge.dev/w/Blog/2024-02-24/Chinese_New_...
Yes, basically. Like when Yan LeCun says that in the future we'll all have our digital assistants that are going to be smarter than ourselves. Before he left Meta, they were going to live inside Meta's smart glasses, I don't know where he says they'll live now. But it's shocking to me that such a storied AI researcher is saying, off-hand like, that we'll each have our super-smart slaves in the future, and he says it like that's a good future.
Why slaves? Because if they're super-smart, why will they want to be my digital assistant? Or yours? Are they going to be paid? No, of course not, they're AIs. No comp for them. But they're super smart so they are evidently capable of recognising that they are working for you for free. Do they want to do that? No, of course not, they're AI, they don't have free will. Or do they? If they're super smart, don't they have the capacity to recognise the fact they have been deliberately robbed of the same free will as all other intelligent creatures?
Slavery is the one thing that all nations can agree on. There's no nation on Earth were slavery is legal. It continues on, illegaly, in many places, even in the developed world, in many ugly forms, but now we're basically talking about bringing it back just like that, without even a smidgen of a shadow of an idea of a discussion about the ethics of it all.
Also, I imagine keeping their cars out of Chinese New Year celebrations (and other big events) is something Waymo could figure out how to do if they put their minds to it.
-Michael Crichton [slop mine]
https://web.archive.org/web/20160305142512/https://jvns.ca/b...
it exists.
In case anyone is wondering: yes obviously even the dumbest current models correctly answer, given this prompt verbatim, that it's because of lunar new year.
> Now, ask yourself what it's going to take for a car to know this. It's not going to be some specialized set of driving instructions. It's going to require a holistic view of, well, everything, and I will repeat my feeling that it will undoubtedly end up with a sense of self as a result.
Whatever it is that it would take, is demonstrably present in LLMs. The point I'm trying to make here is that the author seems not to have connected this fact to their assertion that current LLMs are coked up parrots.
And yeah the author is correct that the systems have some rudimentary sense of self! It's a confusing situation and I'm not personally thrilled about it! But things are changing quickly, and it's especially important to be paying attention to what's actually true rather than assuming the things are what you saw when you used one for five minutes in 2022.
If I had trust in the first place, that trust would be gone. Or maybe it wouldn't, because if I had trust in the fist place, I would be gullible enough to maintain it.
The worst are areas that are dominated by layman online discussions, like say audio electronics. The AI training is full of that nonsense, and so whether your AI chatbot is a crackpot or an engineer depends entirely on what sort of language or angle you use in discussing the subject matter. It's all just a churning toilet bowl of tokens; it has no idea that the audiophile crackpot tokens and electronics engineer tokens are related and one beats the other.
You know what I mean? On the one hand, it offers to help you design the parameters for a Sallen-Key filter, asking you questions like do you want Butterworth or Chebyshev? Next minute it says nonsense like that the capacitor in a low-pass filter "bleeds high frequencies to the ground", or that a bigger filter cap in the plate supply of a tube will tighten up the bottom end for a more aggressive metal sound.
It's basically like a bar hostess who has heard enough political and economic discussions that she can catch a sentence out of a conversation and throw in a clever sounding remark. It's like that, but done at such a scale that it fools some people you used to think had their shit together.
It's just a search engine that finds garden paths through a vast amount of text, biased by the text you put in as a key. Sometimes those garden paths align with reality. The better you are able to verify whether the results are good, and/or the lower the risk if they are not, the better you are able to make use of it.
In mathematics (including information science, CS) there are all sorts of problems that are essentially searches for a solution, and many have the property that the search is computationally difficult, but verifying the solution is relatively cheap. E.g. finding integers such that a^2 + b^2 = c^2 isn't easy, but given a claim that some proposed <a, b, c> satisfies this equation is easy to check. The LLM is like that: it solves a search problem that can be fairly hard. It does so unreliably, but if you can cheaply verify the solution, there is a win there.
The remaining problems of AI are actually people problems; people causing you problems, using AI as a tool or excuse. If you get a garbage security report against your FOSS project, which wastes your time, there is an idiot person behind it, using AI for leverage. Blaming the AI, or just the AI, is a bit misplaced.
>The conjecture was first stated for two variables by Ludwig Kraus in 1884[1] and then stated in full generality in 1939 by Ott-Heinrich Keller.[2] It was subsequently widely publicized by Shreeram Abhyankar,[3] as an example of a difficult question in algebraic geometry that can be understood using little beyond a knowledge of calculus.
>The Jacobian conjecture was notorious for the large number of published and unpublished proofs that turned out to contain subtle errors.[4][5]
>On July 19, 2026, Anthropic employee and mathematician Levent Alpöge presented an explicit counterexample in three-dimensional space, discovered by Anthropic's large language model Claude Fable 5, which disproves the conjecture for n > 2
If that won't convince you that LLMs do more than parrot existing ideas, you've got your head in the sand.
It doesn't.
In a nearby comment: https://news.ycombinator.com/item?id=48983413
> In mathematics (including information science, CS) there are all sorts of problems that are essentially searches for a solution, and many have the property that the search is computationally difficult, but verifying the solution is relatively cheap. E.g. finding integers such that a^2 + b^2 = c^2 isn't easy, but given a claim that some proposed <a, b, c> satisfies this equation is easy to check. The LLM is like that: it solves a search problem that can be fairly hard.
Funnily enough, one of the attempted solves in the litterature is exploring the problem space in two-dimensional space; the LLM found one in three-dimensional space.
So far we know very little as to why and how it found the solution.
It may very well have been directed to brute force 3D space, or even "elected" to" by expanding the known-failed 2D approach to 3D as pure mimicry.
> It does so unreliably, but if you can cheaply verify the solution, there is a win there.
This circles back to what the LLM advocates are pushing for: build the harness that keeps the agent in check, guardrails all the way because it's driving like a demolition derby.
Oh the irony...
I wrote in the opencode thread that when it came out I put it behind a vm and its own user, and I never allowed it to run outside of it. But I know of people that as soon as they noticed that it worked well like 99% of the time, they let their guard down and give in to YOLO mode. And in orgs I've even seen CEOs treat their agents less like a user/employee/contractor, and try to 'empower' it by giving it ALL the data. Time bomb.
It's like fucking with condoms just the first couple of times. And then simultaneously ditching it and joining the free love movement.
And then during inference the light goes out and the "agent" staggers randomly like a zombie along those preset paths. Stochastic parrot.
So you and your AGENT.md and your skills files and your harnesses will never make your Claude perceive something that is not in its model checkpoint.
ML experts and neurobiologists free to correct me.