Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

0% Positive

Analyzed from 427 words in the discussion.

Trending Topics

#why#literature#data#academic#performance#isn#different#exactly#scientific#authors

Discussion (8 Comments)Read Original on HackerNews

nemomarx•about 2 hours ago
Isn't everything like this? Ask an AI about news and it has to deal with contradictory and poorly labeled accounts of an event. Ask it about mechanics in a game and it has to deal with every version of it, maybe different editions or remakes, etc. What dataset exactly is pure and nicely labeled for correctness? Is it large enough for training?

isn't this why we went to synthetic corpuses anyway

ghostly_s•about 2 hours ago
> Imagining a two-by-two matrix with the axes honest vs. dishonest and right vs. wrong, the scientific literature is splashed haphazardly across all four boxes.

Why should we give an extraordinarily claim like this any credence from the authors of a newly-minted substack who can't be assed to write more than a blurb on the topic, and whose listed credentials amount to a cagey statement that isn't even clear on whether they hold degrees?

cwmoore•about 2 hours ago
How does your appeal to authority impact their claim?
cauch•about 1 hour ago
That's not a "appeal to authority", just that reliable conclusions require people who have dedicated enough thoughts on the subject to avoid basic mistakes.

You have exactly the same thing in software development: a closed source software released by an unknown small team that makes surprising claims and that show clues that the authors don't really understand what they are doing is also not trusted.

pandinus•about 2 hours ago
This just in: fallible Humans (un)knowingly produce unreliable data. LLM training on unreliable data impacts accuracy. Some sources of Human-produced data are measurably more reliable than others.

"Trust the science!"

sebastianconcpt•about 2 hours ago
Then is poisonous to us too.
iLoveOncall•about 2 hours ago
Not as much as LLMs are poisonous to the scientific literature (or, really, any literature).
light_hue_1•about 2 hours ago
Ironic that "Reinvent science" would publish such a trash article. Did they even read the original paper? This is not at all what it says.

https://arxiv.org/pdf/2305.13169

The "poisonous scientific literature" removal is on page 14. Look at the table. Removing academic pieces changes performance by on average 0.44%! You could sneeze and change the performance by that much. You could rerun with a different seed and change the performance by more than that. You could retrain on a different GPU that orders floating point operations differently and change performance by that much. etc. This is meaningless.

Also, the authors explain exactly why this happens! It's on that page even. The academic data hurts a little on datasets which aren't academic. It hurts on common sense reasoning like SocialIQA (Q: "Jordan wanted to tell Tracy a secret, so Jordan leaned towards Tracy. Why did Jordan do this?" A: "Make sure no one else could hear"). Shocking that you can't learn this from academic publications.

This is just substandard blog slop that give science a bad name.

The scary part is: "Dan Recht and Ben Reinhardt trained as scientists and now coach scientists. They do other work but don’t link to it here." As someone who has advised plenty of PhD students I can't find the words to express my disdain at that line and these jokers.

nekusar•about 2 hours ago
This article is barely even legible, and chains unrelated papers in some big scare.

Regardless of slop, it's definitely no/low quality and a bunch of breathless garbage.

Flagging and warning others.