RU version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
50% Positive
Analyzed from 7148 words in the discussion.
Trending Topics
#openai#data#problem#problems#more#llm#model#seems#don#used

Discussion (179 Comments)Read Original on HackerNews
I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years.
I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days. Competition is a hell of a drug, and frontier LLMs aggressively compound that energy.
I work for a startup. We often bring a wooden arcade with us to conferences as a marketing gimmick.
The arcade runs a single side-scrolling video game. You're running from a monster and dodging obstacles. The goal is to survive as long as possible, and your result is measured in meters.
There are always a few competitive guys who spend the entire conference taking turns to play it. And every single time, the same thing happens.
Say the current high score is around 200m. Everybody fails somewhere around that number: 190m, 186m... Maybe someone manages 210m. And the high score moves up at a snail's pace.
Then, a new guy shows up and gets something like 500m on his third try. From their next turn on, everybody easily does 450 or more, even though they were struggling to get past 200 just one turn ago.
What makes it stranger is that the game is dead simple. It's not like the new guy discovered a move that unlocked this capability. And it wasn't a lack of motivation either - they'd all been playing for an hour already. They just started performing better after seeing it was possible. There has to be a name for this phenomenon.
Could be described as some physical form of this effect: https://en.wikipedia.org/wiki/Asch_conformity_experiments
The term would be conformity / normative social influence.
A better example would be mixed boys/girls sports classes in school, where the boys deliberately hold back as to not injure/scare the girls.
It's a pretty obvious and human thing not to go out and completely destroy a much weaker opponent. We're social animals after all.
There also may be an element of energy conservation, there's objectively no need to put in any more effort than necessary. Inefficient.
[1] https://www.nber.org/system/files/working_papers/w20343/w203... [2] https://en.wikipedia.org/wiki/Anchoring_effect
If you read the history of major scientific discoveries, this has been the case for a long time. There are many things that were independently discovered by different people at nearly the same time. Once people know something is solved or solvable, it gets a relentless amount of focus.
I'd say it illustrates well that this last piece, the person celebrated for the achievement, is disproportionally overvalued and the rest of the work they are standing on is disproportionally ignored.
I always like to bring up how many decades of research, how many hundreds of years of entire PhD-theses, how many sleepless nights were used up to generate all the protein structure data that made up the corpus of Protein Data Bank - that was then hovered up by the AlphaFold team, and guess who got the Nobel Prize...
On a tangent, the genius of people like eg Einstein is not so much that he came up with all these things: other people were close, but that he was a singular individual that did all of these discoveries, instead of five different guys all making some breakthrough here or there.
Normally, when you tell a coworker that you're wrapping up a result, unless they're some kind of sociopath, their natural inclination would not be to try to steal it from you.
Would anyone be surprised if major model companies had tagged the accounts of competitor employees for extra tracking? Given the concerns about distillation and bench marking it hardly seems irrational, but how it is used matters quite a lot.
What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?
But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI, shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."
The text for all the Goosebumps books are certainly in the training data and to some small amount influenced the solve. But their contribution was so vanishingly small it would seem absurd to say R L Stein should have recourse for contibuting to the solve.
The equivalent would be taking a (fully offline) LLM and asking it about the ending of one specific Goosebumps book, and it revealing the twist. And although that specific book was (probably) only once in the training data, a high parameter LLM can usually "remember" the twist.
- They threw compute on a problem another team/company was rumored to have solved to see what their secret model could do.
- The texts I read do make it seem like OpenAI wanted to talk and share credit generously.
- Imagine working on a frontier math problem with someone at Anthropic and not only do you use Codex but also through a non-business account that allows training on your data.
- Timeline-wise, if they mainly used GPT 5.6 it's unlikely any meaningful data made it into an model that's being internally validated right now.
[Edit: they said "two of": "On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. .. we launched an effort ... on all open Millennium Prize problems".]
Isn't that exactly what almost everyone would do given that they wanted to see how capable their model is and the tense competition they have with Anthropic right now? Stealing impressive headlines from your competitor is pure gold.
If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt an LLM would be useful at all on their own. Is this the «ultimate ASI test»?
However, how to we know the next field of mathematics isn't a novel combinations of several other sub-fields? That level of mathematics would be indistinguishable from magic to most people and so in their eyes the GPT did something truly inventive.
And an LLM could in theory also stumble upon entirely new ideas: there's randomness in how they generate their reasoning and answers after all.
I suspect that we are seeing a lot of advances coming from the combination of existing but somewhat obscure knowledge coming from LLMs at the moment, because LLMs are really good at this. At least compared to humans.
Even before our AI friends became good, they were already known for having read approximately every paper and every textbook published in any language. You only need to increase intelligence a fairly small amount from there to get to something like the 'convex hull' of human knowledge.
Compare https://slatestarcodex.com/2016/11/17/the-alzheimer-photo/
The gist is that basically whenever anyone comes up with a new method you get a big burst of activity of picking up all the now lower hanging fruit, that was previously out of reach.
> If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)
> If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?
> My new preferred hypothetical for this is:
> If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?
This would work just as well and have plausible deniability.
You'll say, "Don't tell anyone", which they will ignore because they get a rush and perceived status by sharing it. So then they tell someone, along with "Don't tell anyone", etc.
It's a small enough world (both in academic math, one at Anthropic) that you get to OAI in very few hops.
Also it's not like you're looking for a lost pair of car keys: just getting to the level where you can understand a problem well enough to "steal" it takes a huge amount of work. People are going to know if you're at the level where you could be a competitor.
So in general it's pretty safe to talk generally about whatever you're doing.
We know that OpenAI trained on their prompts, plagiarism is incredibly likely. The only thing we don't know is whether or not it was deliberate plagiarism yet
How can openai do this, what is being claimed, at the scale of their entire userbase? If they do this only for particular sessions then how do they sieve through sessions for the good stuff?
How are sessions stored, how are they processed, how much storage and how much compute is used in these pipelines, how economical is it, how fast are the requirements on the storage on the compute growing as userbase grows and generated data grows.
All these questions are far more pertinent than the navier-stokes, but i can only imagine all at openai doubling down on this "very important" mathematical milestone.
AKA everything you send to them (and I bet it's the same for any other lab) will be used, no matter what are the TOS, the law or what they publicly say.
I say that as a mathematician (on paper) who perhaps surprisingly doesn’t give a crap about the problem itself.
That’s obvious, isn’t it? Just like you wouldn’t upload your confidential documents to an online spellchecker, or your proprietary code to an online compiler?
I don’t get it.
To me it looked absurd. But I was almost the smallest cog in the corporate structure, so I never asked what were the reasons for all these decisions.
This is something I feared would happen and where open source would be left behind. Maybe they can do something with crowd-sourcing computational power from volunteirs. They were after all able to get Leela Chess Zero to be comparable to AlphaZero by training from volunteer processing power but it seems to me we live in a world now where the best models keep their stuff closed.
In imagine generation too. I'm not sure how well Stable Diffusion can compete in following instructions with all those advanced models that are kept secret.
Nobody has that power. Certainly not mathematicians.
Wait until you learn how many other SaaS and web 2.0 and cloud based things are also run by assholes.
You know this metaphor? https://en.wikipedia.org/wiki/Turtles_all_the_way_down
But instead of turtles, it's assholes.
But more seriously, no, none of what I wrote above is an attempt to excuse or play down the specific role of assholes in large AI companies.
Does google docs own the content of docs you make with it? Does Apple claim ownership of discoveries made using their tools?
Why is all of the world dependent on tech an ever more hostile US? Same answer.
I wish that were true, but I live in the United States and it is 2026.
The President of the United States rug-pulls memecoin crypto and regularly pardons people like Paul Walczak (who was convicted of massive payroll fraud) in exchange for large donations.
I wouldn't make any assumptions about what is considered theft anymore, at least not when it is being committed by people who have enough money to be above the law.
"Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy."
But apparently there is also an entire completely different route "Do not train on my data"?
Does this mean that before I submitted the "Do not train on my data" request, my data was used for training in spite of "Improve the model for everyone" being turned off?
We are getting to facebook/meta-levels of privacy settings obfuscation.
That in itself is a dark, dark pattern. There should at the very least be explicit warnings for users who have checked “do not train”; or they should not be presented with such dialogs.
Maybe next step is to filter your input client side through an unknown number of obfuscators where you ask LLMs to rephrase your question (onion router idea) such that no single provider can be certain that this is human input and not some slop feedback loop.
OpenAI can easily identify these outstanding human behind their accounts. Human in OpenAI constantly check their logs for breakthrough. When they find something interesting, they brute force the result using their massive computing power.
No LLM is even needed.
We know that LLMs are trained to recall relevant info. We know AI vendors are using user transcripts to train models. Two plus two equals four, right? I mean, an LLM that failed to recall the transcripts of those researchers would be a bad model.
I thought this was an ebay thing for people with too much free money, but it seems a bit larger.
Move 37 comes to mind.
2) Because if we accept the facts ChatGPT only came up with its "new" idea after being told exactly what the new idea was by a mathematician (OpenAI doesn't dispute this btw). And OpenAIs story comes down to the usual "We didn't look at it, trust me bro".
3) And, probably, the researchers were likely stopped by token limits, and that's the only reason they were slower than OpenAI themselves, which is very, very unfair.
4) OpenAI's story "smells" (like so many AI stories lately). Supposedly the company's team asked ChatGPT about solving millennium problems, and out of all millennium problems it just happens to pick the one where a solution can be found in its chat logs?
5) Yet again it would be in good taste for these AI companies to just give this to the researchers (no shortage of difficult unsolved math problems, so if AI can solve them all, just find another one). But instead, yet again they're fighting about it.
That 100x step up from using 100 agents for Euler to 100_000 agents for Navier-Stokes, in a single day seems a bit sus.
I also think this is not ethical behavior. This is at least academic dishonesty, kind of a plagiarism or intellectual theft.
That’s why they wanted to credit Tristan and to give $1M award to him. But again, they acted unethically in that process as well. They wanted him to remove Levent (Anthropic affiliation) as co-author and threatened Tristan to “end his career”. Their behavior is actually telling, their work was not completely independent from Tristan&Levent’s unpublished work.
The most direct line from problem to proof is OpenAI building off of conversations the mathematicians had with their AI.
It is able to contribute code, but maybe not good code.
It's the same in math: it's able to solve problems, but not necessarily in a good way with a human readable code.
Math papers are a lot like software:
- theorems are like API
- lemmata like internal/private function API
- definitions are like types
- the proofs are the implementation
The proofs of ChatGPT are not necessarily readable or maintainable.
I think you should not say that. Buckmaster did only state his version of events and was very clear on that he did not make any accusations at all.
To quote from his statement pdf:
> I am not accusing anyone of anything.
Showing of the model's capabilities - okay, but it's not like it solved the problem on its own, and apparently not particularly efficient. Are there practical applications that justify the investment?
There are two benefits though.
One is recognition. Cred. The PhD candidate gets to put a ", Ph.D." behind their name, opening doors to future academic employment or other endeavors where people value titles. The AI lab gets to say their tech solved sth that humanity wanted bad for a long time. Both cases with substantial financial upside (higher income for Mr. PhD and higher company valuation for the AI lab).
The other one is that this is how scientific progress works. $1M or not. That number was just a PR campaign by the math community to point to some goals. It's clear that it would cost more than $1M to get there.
People talk about the "compute" but what about the "storage"? Is storage exponentially greater, or soon to be, than the compute? Is the storage going to slow down growing to some constant rate, i.e. all people on earth using chatgpt, or no, on the contrary, it will keep growing?
If there were any shady business, I do not condone it, but technologically we are not there yet for said shady business to happen.
I agree that they have pipelines for what you are describing but how effective they are at scale and at focusing is the question.
(This was a joke. I value Simon's role in the community.)
What I cannot reconcile is the timeline and the concern in this specific case.
I don’t think training pipelines are anything close to the level of continuous training needed to incorporate Aug 15th ideas into a model that generates a breakthrough early Sept. Either OpenAI nakedly had someone with mathematical understanding dig into a specific user’s chats (a massive red flag) or this really is poor handling of a more classic parallel discovery situation (with one party clearly having worked on it longer)
The amount of millions available to do those kind of research cases is practically unlimited.
There are an unlimited amount of cases to work on.
What a huge development would this give to both humans and the world in general. Because in the end better understanding gives new options.
It's deeply interesting that those things now get a concrete economical price which seems to be viable to extrapolate. The enormous additional "production" of knowledge will inherently increase the speed of all pieces of research and development.
Taken into account that it's used wisely, the risks with a strong force are always huge as well.
A) That we're at the point where SOTA models can, on their own or guided, be used to solve such monumental problems
B) That EVEN if they exist, they're still so cost prohibitive that they are completely out of range for pretty much everyone. Yes, yes, if the costs drop like a stone the hoi polloi can access this power in a year or two - but fundamentally it will divide cutting edge resource into two groups: Those with money, and those without.
That sort of latency, in turn, could lead to some feedback loop where research centers / groups that break barriers get more resources, and those who do not, are starved of resources. This sort of stratification can seriously lead to more centralized research. Do we want a future where only the chosen few get to make progress? For no other reason than that they are the ones with enough resources to spend on the required compute.
Basically, it would be a proof that all the REALLY hard (combinatorial) problems out there, have a much simpler solution, if we were able to find it.
This imbroglio looks like a point in the journey for a firm that is realizing it is not going to make money selling shovels during a gold rush, and that it has more to gain from just… being vertically integrated across an industry.
Open models have definitely hampered the ability to sell tokens at a premium, so mass market adoption is impossible.
But then take someone like Jane Street, for example. They self report making $30bn leveraging LLMs. It’s a defensible assumption that they are making profit on it.
Perhaps it’s more profitable for OpenAI/Anthropic to build their own funds. They have the capital, compute, they can afford to recruit teams and buy any IP/Data required.
This isn’t a fully fleshed out argument, but it is the first time it feels like the winds are changing.
Them being assholes, trying to exclude an author just because he worked at Anthropic, shows the kind of culture within (that part of) their organization. The focus wasn't on supporting academics or expanding research. It was on getting great marketing.
If they had to burn millions of dollars solving a problem _that they thought was already being solved_ to do so, they'd do it.
[1] https://leodemoura.github.io/blog/2026-8-1-postmortem-for-ke...
Edit: at least ~600,000 lines
https://stanfordtechreview.com/articles/openai-buckmaster-na...
[1]: https://www.sqlai.ai
But I don't think openai will bother to release a competitor, the real threat is that anyone with a decent LLM and a harness to try a few queries will land at the same or a better query within minutes.
If this doesn't show you that the fantastical claims of their LLM's ability to solve problems on their own are bullshit, I don't know what will.
It's pretty clear all of this is a marketing effort only, and they pass human results as LLM findings.
We already know that the supposedly industry-changing Mythos and Fable results were actually complete BS and they're just your run of the mill model. There's nothing at all to suggest this is any different, and once we get this "unreleased model" (aka bob from the math department), we'll see it was all lies again.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
I mean... yeah? What do you expect? What else can they say? How could you prove a negative in this case? I do not want to comment on specific OAI employee chat messages, but on the actual OAI discovery here.
> If OAI is using chat data from researchers to make discoveries, essentially every researcher who chats with an LLM can get scooped
But they make very clear that they do train on this if you do not disable the setting. We can comment on the fact that this is opt out instead of opt in, but this discourse of OAI sniping the solution out of some researchers hands seems to be running on the best case speculation of the researchers having perfectly handled all their chats and discussions with other researchers and the worse case of OAI not having full pipeline control and I think that is an unfair assumption.
So what they are saying is that they don't have zero retention.
OpenAI even wanted to respect that part but now you’re too focused on them being able to leverage the treasure map at the same time instead of the treasure for humanity being found at all, a completely different treasure where you both have to deal with getting paid by the bounty provider anyway
Goofy
That aspect alone would/should be unthinkable to any serious academic. Authorship reflects who did the work and changing it for business competition reasons should be a red flag for multiple different reasons. They include, the sheer tactlessness of treating a major theoretical advancement as a competitive posturing first, the norms of academia second, and all the misunderstandings of the culture of the disciplines culture that people will now suspect are hiding beneath the visible surface (insert topography joke).
Math as a field is fairly unique even in how they list authorship. It was long the norm that authorship to be alphabetical because the idea of first, second, senior etc authorship is harder to define than many other fields.
“The stated rationale for alphabetical order is that it treats co-authorship as intellectually joint work: every listed author’s name carries equal weight, and no one has to negotiate, or be seen to negotiate, over billing. That is a genuine advantage over position-coded conventions, where disputes over who is “first author” are one of the most common sources of authorship conflict in fields that use them” [0]
That norm is changing, slowly, but one option people are pursuing is notable: randomized author order. Their is a perception that alphabetical is too biased…that’s the world OpenAI is stepping into when they make that offer of authorship to one scholar with a demand that he exclude his partner.
I can’t speak to the facts of anything else in this, but if a grad student came to me and said someone made them that offer, I would tell them to run and if they were brave report it.
[0] a to the point lay description of the history of math authorship can be found here: https://casrai.org/guides/mathematics-alphabetical-authorshi...
Kind of a tangent, but in some fields authorship is actually more of a formal rank: once you're in the group you appear on every paper. A good fraction of the "authors" won't even know that a paper is being published in their name, and the vast majority haven't read a word of it.
In these fields it can be quite a challenge to not appear on a paper.
https://xcancel.com/SebastienBubeck/status/20973794116915163...