ES version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
43% Positive
Analyzed from 5905 words in the discussion.
Trending Topics
#don#more#wrong#study#llm#answer#textbook#advice#answers#questions

Discussion (203 Comments)Read Original on HackerNews
This study gave people access to an LLM that the researchers knew would give incorrect answers to certain questions, and then quizzed people on those questions, with the option to not respond to a given question if they are unsure about the answer.
This is akin to giving someone a textbook on an obscure subject that has certain factual errors, letting them know they can use that textbook in a quiz on that subject, and then quizzing that person on those facts that the textbook gets wrong.
Obviously that person is both more likely to be willing to respond to the question and is more likely to get it wrong!
There are a lot of things I'm very interested in that are specific to modern LLMs and how they affect learning and confidence (sycophancy, cognitive helplessness, etc.).
This study tested none of those. Its experimental setup is not very different than simply substituting the LLM with a textbook with errors.
A better headline would read, "Very inaccurate AI made people less accurate" but this would make people naturally ask, "what about a reasonably accurate AI?".
It's not about the LLM, it's about whether people will critically evaluate what it spits out.
"The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions."
So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct.
The way the study is organized is like having people hear advice from a doctor who answers questions incorrectly almost without exception, then reporting that people who listen to doctors are 3x less accurate. But that would be an incorrect conclusion because doctors are not wrong almost without exception.
If the question is "how inaccurate does AI advice make people?", then the accuracy of the AI is necessarily a parameter of the answer.
Its about whether people will critically evaluate any information they are given. It has nothing to do with LLMs.
I'll also add that even in these simple experimental conditions, I'd bet that having access to a textbook wouldn't have nearly as much of an effect, for a very simple reason: looking up an answer in a textbook is a lot more work than asking an LLM. So when you don't know and aren't forced to answer, I'd bet it's a lot less likely you'd spend the time to look up the answer in the text book. Even more so if the textbook had "this may contain wrong answers!" printed on the cover, like the AIs do.
Anyone trusting AI as the single authoritative source of information is stupid - but this follows from the fact that trusting anyone as a "singular point" as a source of information is stupid. You corroborate, you intervene on the world to test your mental model, you discuss with other people. That's what learning is. I've never learned from start to back to a textbook before as the single source of information (besides one philosophy of science textbook; in which I spent a month digging around adjacent fields, and then it just so happened that that one textbook synthesized every piece of information I looked up, and it was mostly a consolidating review).
If your study pre-supposes certain courses of action and artificially constrains the action space for the sake of "reproducibility", you may get a result, and a "scientifically rigorous one". But it's not going to say anything about reality in any meaningful way. While anecdotes and the complexity of real life isn't "science" (in that it's a controlled, repeatable, interventional experiment that's subject to a community of critics who want to hold you up to standards of rigor), there's far more truth in how people actually proceed and engage with these tools.
The study doesn't show that at all. It didn't test actual AI.
They could have tested a cohort of subjects with access to actual ChatGPT. Ask yourself why they didn't.
Is that what you're selling us?
So in 18 months, we'll just rinse and repeat?
> The researchers used Step 3.5 Flash, a model that was usually wrong on these questions, precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool.
(emphasis mine)
I agree there are important differences in how textbooks and LLMs are used in real life. This study didn't explore that at all. It used a setup that essentially elided the difference between the two.
This is why I think it's a bad study. It didn't measure anything of the essential differences of how people use LLMs.
What open book quizzes allow you to leave all answers blank with no penalty? An open book quiz is very different from the experimental setup tested here.
The number of people using LLMs must dwarf the number of people using textbooks for any reason.
Feels like this differs wildly depending on who you consider "people" to be. The average person on the street? Definitely just parrots stuff they've read somewhere, not even a "textbook". A group of software developers used to parsing semi-true information? Probably they'd get it right, yeah.
Or a teacher they met in childhood who taught them everything they know, right or wrong.
They are the first to parrot what was spewed from llm and previous even what was found on 4chan.
How often do managers just regurgitate ai advice rather than consulting their experts? How often does a person question an expert because the ai said so?
Naturally, the ai will be right some of the time - but it’s really hard to correct for the times the ai is wrong.
We know people are going to defer. The solution here is to make AI (including the free tiers) more reliable.
The study is OK. The article (and the original headline that came with it) is pretty bad because it claims things that the study doesn't. And I guess it is ironic that the TNW article looks 100% AI-generated.
I'd wager you get similar results if you gave people a version of Google search that purposely gave you bad results. Like, it's framed as an assistant / lookup tool - is it so surprising that people tend to trust it more? Especially since the participants are likely used to using full-powered models and the researchers give them a purposely gimped one (lol)
People are acting rationally when given AI tools to lookup information, their first consumer use case was as a super-powered Google Search
Yes and no.
I think most people would agree that Wikipedia is, on the whole, a pretty great first resource on anything. It tries to be factual and accurate.
Most people would also agree that Wikipedia can be wrong or manipulated and should never be used for an authoritative source.
And then somehow a computer barfing up words distilled from magic internet concentrate is absolutely trustworthy?
I don't get it.
That attitude seems to have gone away for some reason
But in reality when google gives you the wrong answer, you at least have some signals you can use to infer confidence. For example, the number of results, whether the sources are trustworthy, etc.
AI at best tucks that away in a footnote and discourages further critical thinking.
I hadn't even considered people might evaluate knowledge that way. That's legit horrific lol.
"What's the literal odds this info is wrong" vs "is this answer consistent with everything else I know, and if not, what other info would I need to change my mind"
That strikes me as an incredibly appropriate test because LLM’s are unreliable with factual statements. People need to be able to understand that and not treat them like textbooks which are basically 99.9% accurate (let’s please not bicker over the 99.9%. It’s close enough. A major textbook is safe to treat as accurate, an LLM is not).
"This study tested none of those" So the study is bunk because it didn't test your favorite LLM flaws?
Except LLM isn't a textbook, people know that but believe it nonetheless.
The point is that LLMs aren’t right, and the people who took the test were probably reminded of that.
Would people have trusted the textbook you’re mentioning if there was a big red warning on each page that said “this book may contain errors”.
The willingness to trust AI even though it may be wrong and even though there’s money on the line is intersting enough as a study imo
Relentless grounding is my personal solution for my AI epistemic crisis, but it's an expensive solution in terms of time and effort, and triage is hard too.
So for example, you might have a company chatbot that responds with a summary of a Confluence page it was trained on and a link to the source.
Doesn't eliminate hallucinations, but it does reduce them and gives you a chance to verify.
what differentiates you? you could tell me "type xyz into chatgpt" or even share a structured document link from their site. but when you're copy/pasting, then you're basically useless in the equation.
That’s crazy, it seems like such an obviously fair policy to me. They rarely even read their own generated text. I don’t understand this mentality
It wasn't as much time as you felt entitled to from them, so you responded by minimizing their contribution and being condescending. That upset them.
The problem here was not the other person.
Yes I could get some of this information off of IMDB or a DIY video but I'm looking for color commentary and regional advice, and also in some of these situations, the people I asked know me. They know the sorts of things I don't need help on and the ones I struggle with. So maybe they can tell me this movie or trick isn't for me and I should try this other thing instead. Or which store I should go to for this.
At the same time, we have historically seen a lot of "conversational bidding" lead to low quality shitposting at best. At worst, we see wildly out of touch echo chambers fueled by bots, guerilla marketing, and political agendas.
You did solve your problem by seeking higher quality people to talk to. I'm just saying there is more to rejection than the other party "not understanding".
In the context of work chat or certain events, I could understand a LMGTFY response as steering conversation out of the mud and saving everyone the headache. It also helps you, the recipient of that response, save face.
They're the ones looking like an asshole while you get another chance to try somewhere else. I've met some of the most sociable people upon realizing their usage of these "jiu-jitsu moves" were not meant to be hostile.
To you the "bid" feels like a genuine attempt to connect, that may be a bit indirect.
To some people it feels more like you are being disingenuous or awkward. They instinctively feel that your intent doesnt match your words and it sets off alarms bells.
The fact that you're on the internet judging the other party negatively for having a different communication style is an indicator of which party was least willing to adapt.
It's usually not the person who gets asked who acts like an asshole, and if you don't don't want to be judged by people for your behavior, start by not openly and loudly judging them for theirs.
No sympathy for this sort of behavior and you're out of line. Or you're the problem. I'll let you decide since I don't know you.
I would encourage anyone considering a reply like this to ask themselves what value they have brought to the table over the OP using AI themselves.
Most of the advice and information subreddits I previously visited collapsed into bad posting even before AI.
The few I still follow are oddly fine.
The key was always having an active moderator who cared and had unbelievable amounts of time to moderate. Once the subreddit gets past a critical mass of junk-posting users, the good users leave and there's no coming back.
Sometimes the mods are the junk posting users.
A bunch of the car subs were modded for a long time by some tow truck driver who'd mod them while sitting around waiting for dispatch. Exactly the kind of guy you want refereeing a dispute between some guy who's lived it and a bunch of lube techs who are quoting textbook best practices.
These people have always existed and gave bad answers, but now they can do it so much faster, and with so much less actual thinking.
The difference is that nowadays people can just skip the research part and copy-paste Llama. No value added.
and why would the person doing that do that? if the person asking the question wanted to know what an llm thought, wouldn't they have just asked the llm for an instant response instead of having to wait for someone else to respond? i'm at a total loss of being able to understand the point. it's obviously different if the response is a bot posting just in case that needs to be explained to a bot.
It goes a long way.
How did people respond to that?
They cautioned that in some cultures, if you asked for directions, people would rather be polite/helpful and give you an answer rather than saying "I don't know" and giving you no answer. (They recommended that you ask more than one person to be more certain you were getting good directions)
In other cultures they were more likely to say "I don't know" at the cost of seeming more impolite.
So, we don't know the intent. :)
That seems to be the key and its risky when the measure becomes the target itself. When you already know the benchmarks, and that drives the definition of success outcome then there is every incentive to just chase them
One big pet peeve of mine on reddit is on any kind of immigration-related subreddit it is absolutely filled with people talking about how it's impossible to migrate and you'll never get a visa. Yet, meanwhile, people do migrate successfully and legally all the time.
With AI, people literally just act as token proxies and are doing it purely for internet points (upvotes, followers, etc)... I guess on some level, they're just providing a service that the "market" implicitly values. But also, in Bob Slydell's voice: "what would you say you do here?"
The internet as a public forum is dead. And given the purpose of something is what it does, this is the true goal of AI. I'm certain blowing up the internet benefits very many wealthy and powerful people.
That and all the stupid questions.
>but instead someone to relay the question to ChatGPT and post the result as if it is their own hard earned knowledge and insight.
>people aren’t just refusing to say “I don’t know” they’re actively seeking out opportunities to pretend they know things.
What? This is a joke right? Advice and information subreddits have been shallow blind leading blind crap like that and have been for at least a decade. The vote mechanism plus local culture rewards shallow "teennager just googled it" type responses (that AI is basically replacing here) that even people who don't know anything can agree on and punishes nuance and serious understanding.
Eventually the people who know their shit make a joke subreddit (which is inherently exclusive because you can't joke about something without knowing about it) and then at some point later people figure out that's where all the smart people are, start asking questions there instead and then it goes to shit in the same way.
AI is absolutely a lateral move here.
Even if AI gets smarter it will still be agreeable, and people will use it to more confidently reinforce thier stupider ideas, especially in areas where they lack the knowledge to know if they are right.
Even if this can be solved, technically, people won't want to use the model that says they are wrong, so they will choose the glib lies that reinforce their beliefs.
Freedown of speech forces people to think they may be wrong, but freedom of association lets people avoid this. AI is going to be internet hug box echo chambers at an unbelievable scale.
It feels like the example was cherry-picked, but the combination of "less accurate" and "more confident" made me immediately think of this disaster. I truly worry that it's going to take civilian deaths / liability before the proper use / place of these new technologies is adopted, but we could probably start with 1) not in safety critical systems, and 2) don't replace your specialists / experts with algorithms.
[1] https://www.engr.psu.edu/ae/thesis/failures/MKP/failures/fai...
But for the people who do actual work with AI, who are the people who pay a lot of money for it, they prefer to be called wrong when they are, because it is better to be called wrong by an AI than being wrong in front of a customer.
It is not an easy problem to fix, because as much as we want an AI to give us the right answer rather than just being agreeable, we still need them to follow orders, and AI that doesn't would be quite useless. So there is a balance to be found. Smarter models tend to do better because they know what is right, compared to smaller models that don't and then assume you are right and hallucinate from there.
I do like the idea of AI personalities ( I actually use them a fair bit for varying opposing perspectives ), but I am not entirely certain people want that.
As someone who rarely uses LLMs because I dont need to, nor does it benefit me - I work on stuff that is original for which LLMs are useless at - Im glad.
I want everyone around me to get dumb as hell. It makes my path in life much easier and more successful.
Closer to you, you don’t want your customers, your coworkers, your friends, and your neighbors to be dumb as hell. Whatever you may gain by being the smartest one on your block is more than offset by the myriad ways in which dumb people make life worse.
We’ve invented a technology that hijacks that part of humanity more effectively than anything ever. Asking people to “not be so stupid” is not going to help.
> Humans are wired to anthropomorphise; we do it with sticks and rocks.
I do not see people anthropomorphize their TVs, books, paintings, cars, houses, rocks, sticks, ... I'm not sure what you're talking about here. I haven't seen someone call a rock or stick by a pronoun, for example.
I've seen people fancifully do it, but it's exaggeration for a cute effect, and works because everyone knows they're exaggerating.
> We’ve invented a technology that hijacks that part of humanity more effectively than anything ever. Asking people to “not be so stupid” is not going to help.
People learn to look at things in different ways. Humans learn, society changes. It's changed quite a bit in recent years; it's much different than a century ago, a century before that, etc.
In a way, it's more worship of technology to say humans are hopelessly, powerlessly dumb, and technology inevitably rules them. The inevitability of technology is of course a self-serving trope of the industry: you are powerless, we are inevitable. How convenient, and short-sighted.
It's not just fanboys doing that, but "official" educational channels. Both an internal training (more like an all-hands knowledge share) and some dumb training session by a guy from Anthropic explicitly, repeatedly recommended users anthropomorphize their chatbots on no uncertain terms. "Treat it as your super-smart coworkers", "I named mine Becky and I go to her with all my work challenges", etc.
More like hotbox of one's own farts, but yeah
Richard Feynman: "I have the advantage of having found out how hard it is to get to really know something, how careful you have to be about checking the experiments, how easy it is to make mistakes and fool yourself..."
https://www.youtube.com/watch?v=tWr39Q9vBgo
I am inclined to believe the effects of the study are real, but not nearly as pronounced as the data. If there were more serious amounts of money on the table, I think common sense would prevail.
Marcoccia, C., et al. AI Advice Suppresses People’s Willingness to Say “I Don’t Know”, Even When the Advice Is Wrong and Accuracy Is Incentivized. PsyArXiv, 15 July 2026, osf.io/preprints/psyarxiv/5y6m4_v1
Id want to know if “AI” makes a material difference vs just having access to the wrong answer. Like someone could be given search access that successfully retrieved wrong answers to questions, would that give the same results. How much do uniquely AI characteristics, like sycophancy or the conversational aspect play into this, vs people just being willing to believe what they read?
TLDR: They studied both cases (Access to a LLM/Chat interface which gives a wrong answer when asked, and access to pregenerated (wrong) answers). Both experiments yielded similar results.
I feel like this might be an interesting method for places where the default response tends to be someone submitting copy-pasta: just give the ai-response and then have community discussion around what it missed.
We're gonna continue to push for scientific based approaches to education not some lunatic's idea of an educational experiment that waste's children's education. We tried this with Bill Gates already.
You can’t say it didn’t warn us, it’s on the tin: “AI might be wrong.”
I get why they used questions where AI models fail, but it also really reduces the value of this study. Nobody is really asking AI the color of a team's uniform in a movie, and if they do and confidently get it wrong, it just doesn't matter at all.
Asking trivial questions also feels like it would affect the rate at which people are willing to confidently say things that are wrong. If you ask me some question of pointless trivia and I ask ChatGPT, I'll probably just repeat the answer because who cares. If you ask me something even mildly important and I ask ChatGPT, I'll either verify the information before I repeat it to you, or I'll qualify that I looked it up with ChatGPT and didn't verify. But some things are just so unimportant that they don't even warrant the disclaimer.
> Researchers found AI advice suppressed judgment suspension from 44% to 3%, accuracy from 27% to 9%, while confidence rose from 30% to 76%. People trusted wrong AI answers.
I didn't see the article mentioned what they were actually asked, but I'm surprised confidence was only ~30%.
> sudy
Intentional?
And people who follow bad advice will get bad results, and people who follow good advice will get good results?
I wonder if this will impact the quality and prevalence of certain AI models in the future.
https://www.youtube.com/watch?v=axOcn--n_lM
https://www.anthropic.com/research/AI-assistance-coding-skil...
We should remember LLM do have legitimate use-cases like search, as we enter the "Trough of disillusionment" in the hype cycle. =3
https://en.wikipedia.org/wiki/Gartner_hype_cycle
There's plenty of other issues with the study, one of the primary ones being that they chose a relatively 'dumb' AI (Step 3.5 Flash), but also, specifically hand-selected wrong answers. People that are used to using competent AIs that are mostly correct would be operating off of experience that suggests they should trust AI. In this case, by design, they shouldn't have, but it's hard to fault the user for that.