Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

50% Positive

Analyzed from 497 words in the discussion.

Trending Topics

#openai#chatgpt#model#used#researchers#data#everyone#models#more#should

Discussion (7 Comments)Read Original on HackerNews

jrflo•about 1 hour ago
I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:

> Improve the model for everyone

> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.

It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

ColinWright•30 minutes ago
I refer you to this:

https://news.ycombinator.com/item?id=49643556

Quoting:

> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."

asimpleusecase•21 minutes ago
Old Facebook trick - likely resetting that box each time the app is updated.
nmfisher•36 minutes ago
There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem".

I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).

fritzo•10 minutes ago
Whoa that's a slippery slope! Next you'll want model runners to cite the data their models were trained on
omnicognate•43 minutes ago
Not unticking a box in settings doesn't constitute consent in my opinion. I'd never put anything I value into ChatGPT anyway, though.
rfgplk•34 minutes ago
Under EU rules it doesn't constitute consent.
spindump8930•33 minutes ago
"Improve the model for everyone" can be implemented in so many ambiguous ways.

https://news.ycombinator.com/item?id=49643513

fithisux•22 minutes ago
Ok, you shut it down, or that is what they make you believe. You give the instruction to shut down, you can't know if it has been applied.
ColinWright•about 1 hour ago
techblueberry•about 1 hour ago
But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?
Robotbeat•about 1 hour ago
Neither? Competitive academic researchers are susceptible to exaggeration and self-aggrandizing, and CEOs are that and also mostly psychopaths. I tend to think there isn’t systematic spying on researchers looking for breakthroughs. A lot of people are looking for the same things using similar approaches.
techblueberry•43 minutes ago
I mean the accusation is that they were using private ChatGPT conversations. Given the extent to of the gold rush and the long history of Silicon Valley stealing ideas, and arguing it’s not immoral, It almost seems like your making the exceptional claim that this is the one time where Silicon Valley didn’t use information that was at their disposal.

Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.

HarHarVeryFunny•2 minutes ago
It seems that in this case OpenAI are suggesting that the researchers whose work they scooped were using OpenAI models with an account setting that allowed OpenAI to train on anonymized prompts.

It seems that Buckmaster and Levant (who is an Anthropic employee) were rather naive in the amount of trust they had in OpenAI, with Buckmaster going so far as to contact OpenAI's Sébastien Bubeck to discuss what they were working on and clarify that contrary to rumor this was a private collab.

bigstrat2003•21 minutes ago
If Sam Altman tells you the sky is blue, you should double check. I certainly hope nobody believes him when he claims controversial things from which he stands to benefit.
hn1rig3rak•about 1 hour ago
the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.
spindump8930•41 minutes ago
The canary string was more about inadvertent scraping or analysis in other papers. Not direct training on user data. And the use of BB has eroded quite a bit, with BB-Hard or other variants being typically used.