ES version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
18% Positive
Analyzed from 789 words in the discussion.
Trending Topics
#scientific#literature#claim#authors#data#llms#performance#training#why#claims

Discussion (10 Comments)Read Original on HackerNews
isn't this why we went to synthetic corpuses anyway
Why should we give an extraordinarily claim like this any credence from the authors of a newly-minted substack who can't be assed to write more than a blurb on the topic, and whose listed credentials amount to a cagey statement that isn't even clear on whether they hold degrees?
You have exactly the same thing in software development: a closed source software released by an unknown small team that makes surprising claims and that show clues that the authors don't really understand what they are doing is also not trusted.
"Trust the science!"
https://arxiv.org/pdf/2305.13169
The "poisonous scientific literature" removal is on page 14. Look at the table. Removing academic pieces changes performance by on average 0.44%! You could sneeze and change the performance by that much. You could rerun with a different seed and change the performance by more than that. You could retrain on a different GPU that orders floating point operations differently and change performance by that much. etc. This is meaningless.
Also, the authors explain exactly why this happens! It's on that page even. The academic data hurts a little on datasets which aren't academic. It hurts on common sense reasoning like SocialIQA (Q: "Jordan wanted to tell Tracy a secret, so Jordan leaned towards Tracy. Why did Jordan do this?" A: "Make sure no one else could hear"). Shocking that you can't learn this from academic publications.
This is just substandard blog slop that give science a bad name.
The scary part is: "Dan Recht and Ben Reinhardt trained as scientists and now coach scientists. They do other work but don’t link to it here." As someone who has advised plenty of PhD students I can't find the words to express my disdain at that line and these jokers.
1. The claim of the authors is that the scientific literature is on average so dishonest and/or wrong that access to it (pre or post training, I assume?) actually harms LLM accuracy. Hopefully we can all agree this is a claim so extraordinary that would require some very uniquely powerful evidence. After all, we know that LLMs also train on very large corpuses of text with varying degrees of both accuracy and dishonesty - not the least, most of the internet! So the claim of the authors must be that the scientific literature is so bad that it is actually uniquely bad for LLMs, like worse than the general internet.
2. They provide one citation of evidence supporting their claim, which is this paper [0]. If you actually read this paper (which is mostly not about this actual question, but related questions about prefiltering), it provides no evidence for their claims. For example, in Figure 5, removing PubMed, aka the entire biological scientific literature, has by far the strongest negative impact on evaluation of Biomedical questions. It even has a strong negative impact on evaluation of questions in the "Common sense" category! To quote:
"Performance degrades when we remove domains with close alignment between the pre- training and downstream data sources: removing PubMed hurts the BioMed QA evaluations"
3. This post is maybe what you would get if you prompted an LLM: "please provide a citation for the claim that the scientific literature harms LLMs". It's really quite worrying that we're looking at AI generated propaganda designed to convince the reader that the scientific literature is worse than useless.
4. As a scientist, the primary use I want from LLMs outside of code generation is to be a fast and comprehensive search engine. I want to see the papers behind their claims and evaluate them. Literally the only time they are useful in scientific research is when they provide citations for the claims, so I can read the papers and evaluate.
5. Somehow we have gotten to the point where many people, especially people in tech, believe that the scientific literature is mostly junk or mostly useless. This contrasts so starkly with the current rapid pace of genuine, meaningful, society-impacting scientific progress in pretty much every major field. How we have gotten to the point where this myth is so pervasive, I do not know. Is it really spurred just by a few recurring news stories about reproducibility? By just interpersonal bitterness and feelings of anti-'elite' sentiment? (how on Earth your local state university climate scientist is a member of the 'elite' but software engineers making $500k a year or X influencers with millions of followers are not will always simply be beyond me).
[0]. https://aclanthology.org/2024.naacl-long.179/
Regardless of slop, it's definitely no/low quality and a bunch of breathless garbage.
Flagging and warning others.