RU version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
66% Positive
Analyzed from 2785 words in the discussion.
Trending Topics
#openai#data#training#credit#without#mathematicians#user#ideas#why#humanity

Discussion (70 Comments)Read Original on HackerNews
https://mathstodon.xyz/@tao/117237320796901560
Especially in recent years the mathematics community has worked very much in good faith, and a lot of effort is spent trying to give appropriate credit for ideas. Even when ideas are discovered in parallel if it turns out that previous work contained the same essential idea it is by and far regarded as best practice to give priority in this case. The point, in good academic practice, is to maintain the health of the practice at large.
As Terence Tao explains in that post, the goal of mathematics is not only to solve big problems. And, even if one were very single-mindedly focused on solving big problems, it is still (in the long term) better to maintain the health of the community at large so that problems which are out of reach at the moment may be in reach again in the future. Good academic practice is one part of this culture.
The math community was one of the first to be affected because RLVR makes Math an easier target for AI. We witness AI eating the math community.
Software community was also affected due to similar and other reasons. All other professional communities will face a similar challenge in the near future.
If I were a mathematician I would not my unpublished work to go into the hands of a competitor.
If I were a lawyer I wouldn't want private details of my defense to be made available to the prosecution. Anonymous or otherwise.
I wouldn't want the plot to an unreleased book to be suggested to another author.
Why is it OK if openAI does it?
If my friend showed me some unpublished work (Say 90% of the hardest work toward a significant proof) if I then go finish the remaining 10% and publish the whole thing as mine it isn’t ‘building on the shoulders of giants’ it’s plagiarism/theft.
If he’d published that work first and I took it and found another proof and credited him for his work via citation then that would be fine.
But if they are saying this contaminated the model's training data with knowledge of their ideas, who's to say the model they were developing their research ideas with wasn't already contaminated through prior discussions with other researchers about the same topics?
So if contamination is proved, or cannot be disproved, and if these researchers want OpenAI to relinquish its claim to have solved these problems independently, then it would seem they also have to give up their claim to have solved them independently?
2. If someone did come out and claim that their ideas were used without proper attribution in Tristan's work, then of course, that deserves consideration.
3. What you're suggesting seems purely hypothetical. At present, there is nobody claiming that Tristan's work is "contaminated"
4. Tristan was very willing in his initial statement to give credit to the people who developed the ideas.
In the best case scenario, the mathematicians were standing on the shoulders of the extensive training data from sources that aren't being credited and they may not even have had access to.
I have nothing against these mathematicians because I don't know them. And given that, there's no reason to trust their word any more than OpenAI's.
It's possible the work they were doing, even if related, was a dead end and immaterial to OpenAI's findings. Or maybe they are right and OpenAI stole their work. Who knows the truth right now?
In other words, we need more evidence before making accusations.
However, the mathematicians could easily declare whether they had the toggle on or off. Yet curiously, they will not say!
It was very weird on it's own almost like the setting didn't matter. Or the researcher was being dishonest with what they shared to The Verge
I don't know what format they use for storage, but Iceberg would be a reasonable choice. A date in iceberg format is 4 bytes[1]. I checked postgres as well as a reference point. It also uses 4 bytes for a date, so whatever they use it's going to be about that.
Current world population is just shy of 8.3 Billion people [2].
4 bytes times 8.3 billion people gives 30.92 GiB. [3] OpenAI's training data will be in the petabyte range at least.
[1] https://iceberg.apache.org/spec/#schema-evolution
[2] https://worldpopulationreview.com/ and elsewhere, say census.gov if you want a US source https://www.census.gov/popclock/world
[3] https://www.wolframalpha.com/input?i=4+bytes+*+8.3+billion+i...
I'm not criticizing them, but I hope this race towards the first-best result or AGI doesn't blind them to making good decisions such as not using their users' data without consent.
Imagine what a genuinely openness-focused organization of this sort could be. Even if we imagined a commercial half, we could imagine a foundation with mass-membership, perhaps with a membership fee equal to 1/2 the typical personal subscription and functioning to set the direction, elect the board, etc., and then a commercial half which might be rough, tricky, deceptive, making deals with anybody.
I think I'd have been fine with the commercial half being a bit of a monster, as long as I'm part of the members and we decide what sort of board it gets and there's a clear "this is basically controlled by the public" and if I were part of a club of this sort, I would, like you absolutely fill up a directory with texts and computer programs and careful annotations to aid training.
and they could have had it. It could have been easy to make an organization like this. I think you still can. An international AI club, the members vote on what sort of training material may be supplied and for what intents, create some committees to review quality, and then everyone starts making their little games and RL environments and annotated stories and programs that ordinary LLMs misunderstand, and then they get together and fine-tune something, and if that works well they then get some staff and better training infrastructure and end up with a commercial half.
But later we found that algorithm and math cannot be patented.
OpenAI, as with all AI companies, openly admits that it trains on user data unless the user opts out. But the mathematicians have not said whether or not they opted out.
>Did the transcripts of any of the agents include a tool call whose result including user data?
They have already explicitly denied this.
What the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question!
Most people do not understand that the main reason for the subscriptions is to give OpenAI and Anthropic the priceless, unique data that shows how the models are used, what people are building, how they are building, which solutions they consider OK, which they consider bad -- they purchase this data with cheap tokens. This is their only moat, really. If some really proprietary IP gets swept in the training data set its not really OpenAI's fault -- its the researchers'. Have something secretive? Dont fricking paste this into chatgpt. Duh!
(I'd definitely not think OpenAI/Anthropic ignore the opt outs, or ZDRs. All it would take is one whistleblower to get them into terminal troubles. And why would they do it? They are not in the business of scooping unique IP -- they are in the business of understanding how AI is used across a variety of mundane, day to day work of individuals and companies. Useless math problem is good (or bad, as in this case) PR, but otherwise entirely worthless for the labs.
[0] https://mathstodon.xyz/@andreasthom/117240535270608201
[1] https://news.ycombinator.com/item?id=49638353
What I was referring to is the fact that neither Levent Alpöge nor Tristan Buckmaster will answer this question.
What I was referring to is the fact that neither Levent Alpöge nor Tristan Buckmaster will answer this question.
Dude makes up like 50% of the replies here.
"Just use some polynomial" would be an idea so generic that it wouldn't merit attribution. "Just use a cyclotomic polynomial of degree 24" is a lot more concrete.
(I have no idea where the exact limit lies and I have no idea how precise or vague were the ideas in those discussions.)
Would this data moved through a hack to ChatGPT, this would be another thing, but like this. No pity at all.
Re-using advances done by others is necessary in maths.
This is how hard sciences do progress.
On my open source projects, there is always a big "WIP" messy phase which I don't "really" publish... because it is messy.
in most jurisdictions it'd constitute a crime.
So I see someone dedicating a sizeable portion of their life on a discovery and not being credited is the problem here. Otherwise it would be published in a journal anyway.
When your work is unimportant enough for you to not setup a local AI, just shut up.
Just like their "oops we hacked hugging face".
Everything OpenAI does is to drive hype. They aren't honest or a good actor but they sure have good hype.
2. There was a $1,000,000 reward
This kind of bribing attempt doesn't seem to come from the "good side"
Anyone with a decent amount of social power is aware and skilled at maintaining these conditions whether consciously or not.
The praxes are non-trivial: awareness, bravery, flexibility, and seflessness.
This could be summarized as "unhealthy competition."
And lying about whether the proof was done by mathematicians or OpenAI is bad humanity, so.