DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
68% Positive
Analyzed from 3584 words in the discussion.
Trending Topics
#more#don#models#model#down#better#should#everyone#years#agree

Discussion (90 Comments)Read Original on HackerNews
My new suspicion is now that they didn’t got drastically smarter, but they got trained on the user input on the previous generations. I don’t think anymore we had a massive intelligence jump. It just seems like they know more edge cases. Therefore I see this as marketing.
Happy to discuss.
Yes - that’s one of the ways in which they got smarter. More data more better.
I presume also lessons learned re training procedures, algorithmic advances, but data is a primary lever.
I think this happened around Opus 4.5 or 4.5 and the same for GPT 5.4.
The more I use those models, harnesses, techniques for guidance etc etc. The more I land in going back to writing software by hand again. Maybe not all of it, but at least the crucial parts + foundations.
Not among the public-facing models, they're indeed stagnating. However, the development of models for military use won't be slowed down, that much is certain.
> Therefore I see this as marketing.
It's some marketing but mostly politics, it's an attempt to discourage others from developing AI countermeasures to what is being developed in secret. And to fulfill the backstage agreements which aren't worth the paper they aren't written on.
It is not a universal opinion at all that general coding ability has plateaued.
Do you think the majority of developers agree with you? If that was the case wouldn't there be much less disruption of the SWE industry?
I can't remember the last time I even opened VSCode to even check something let alone to fix it.
When I have a clean codebase it’s super powerful and faster than I am. Then I start to use it more, more sessions and longer tasks less checking in between.
It kind of works but later I’m in a deadlock where every change introduces new bugs or takes ages. This might be for a lot of reasons for example me going to fast, me losing mental model, me explaining it wrongly.
However when I then start checking the code it’s all spaghetti like frankly the spaghetti Astra produces I’ve never seen before. Processes that should be simple stretch over 11 files with weird wrappers and abstractions and I need a whole day to entangle it.
These are ai assisted user workflows that Im working on in this case.
I just have the feeling no matter what AI just always expands it. And expansions hinders agility and sometimes you need that.
Instead of burning tokens doing intellectually impressive “hacks”, or sandbox escapes from which I have concerns could contain run of the mill malware, why not focus on showing people what you can fix?
Or, make software people actually use every day better.
Or, donate efforts towards medical research.
I don’t like equating AI to nuclear energy, but, it’s a lot like continually showing people how large of an explosion you can create instead of showing them how many houses/hospitals/schools/etc you can power.
Classy
This could be the latter showing its face.
Unlike with nuclear proliferation, there is no heavy industrial base requirement. No time consuming, visible uranium enrichment. All it takes is for someone to buy sufficient amount of compute and try to get it past the RSI gate. Or bribe people with access to model weights to existing frontier - the asymmetry between what it takes to bribe a bunch of geeks vs. what is at stake is staggering.
In almost every technical industry, China is extremely good at fast-copying and relatively mediocre at solving problems that haven't even been posed yet.
Even if they were 95% as good, but 30% of the price, they would win
fear is still the best tool to get masses to agree with whatever plan you hatch for yourself
I, as a human, would greatly prefer just another corporate entity with economic funny business over AI destroying humanity
Sure I’d prefer neither, but what are we talking about here? Nearly everyone building these systems is outspokenly concerned about grave consequences. Existential.
That is what they are claiming. I do not believe them.
> Jesus Dario we get it man, you want clout for the IPO.
Having third party embedded researchers red teaming alignment is OK.
The alternative is having a bureaucratic agency like the US FDA testing and vetting models. The government probably couldn't keep pace with AI research right now, but it's where we'll eventually end up at anyway.
Fable/Astra can generate massive income for years even with no further advances.
Dario, Sam and Elon would not agree to do this if they didn't think alignment is possible in the short term, and the pain incurred by third party evaluators would be limited. Surely, if recursive self improvement works wonders in AI research, it can similarly work wonders in alignment. So this could be a PR stunt.
P-doom not negligible indeed.
i have been told in the past that any idea that starts with "if everyone just..." would likely fail, because everyone would not just.
problem is dario's arguments are highly binary. without a dealmaker, and giving the leverage away at the first time of asking makes it harder to work.
People have been saying this for 3 years now. Eppur si muove...
So, the first set of questions is would 'exponentially better LLMs' dramatically increase the probability of any of the above, or domains that Dario is not citing? That assumes that exponential improvements will happen if there is not 'pacing'.
IF answers to above are 'yes', then we need to question if 'pacing' is viable. To use a different domain, regulating 95% of vehicles to a max speed would likely save 100s of 1000s of lives, but is not perceived to be viable. In other examples, regulation has unintended consequences in the opposite direction (e.g. some 'rent control' efforts and arguably some drug/alcohol laws).
OpenAI considers slowing advanced AI development, Sam Altman tells employees
Link: https://www.bloomberg.com/news/articles/2026-09-11/openai-is...
HN Post: https://news.ycombinator.com/item?id=49652270
What if instead of building one big scary AI god, we built an ecosystem of extremely effective, efficient and predictable, reliable tools?
Tools that are specialized for particular purposes could potentially be safer, more efficient, more reliable and effective, and even more profitable for their makers. We currently put science, engineering, general question answering, paper-writing, "smart search", and chatting with a fantasy character all into the same token-prediction platform. We struggle with hallucinations when trying to make fact-based decisions using the same tool that our neighbor might be using for creating writing prompts (or at least was trained in part on fanfic).
We pushed people to integrate giant generalist models into their workflows, tie in with tools, etc. And then a coworker can have a bot write slack updates that summarize progress on tickets from your teams project, at the cost of giving an untrustworthy agent access to a bunch of internal systems, and a small risk that it will do something crazy. And when it works, that's worth _something_ but probably our team would pay for that ability at a different price point than the models we use to build products.
If I could do programming with a faster, specialized model that was trained not just to complete program text-tokens but on tuples of program text, compiler IR, program traces, etc, so it had a deep and explicit understanding of how the program would build and execute, and where the model was closely integrated with the language toolchain, I might be more productive and be willing to pay more than for a generalist model. I don't _need_ my coding model to be able to role play, or be able to potentially engineer a super-pathogen, and if the size and latency of my coding model could be lowered by entirely cutting out that possibility, everyone can be better off.
Maybe bioscience applications are super valuable, in which case someone should build them. But does the model that suggests CRISPR edits to a model bacterium need to also know about computer security, and should it have scifi novels in its training data? Does it even need to be capable of producing unconstrained tokens, or should it be limited to producing in some relevant domain-specific language? And would labs be willing to pay more for a model that was specialized, and by construction unable to try to break out of its sandbox and post their experiment protocols to an obscure german wiki?
https://darioamodei.com/post/we-must-pace-the-frontier#top
Consider for example the question "What caused the French Revolution?" Many different answers could be given, all technically correct. What gets emphasized is where the ideology lives.
Somehow I think if these companies were held financially responsible for what their AI did, we would get alignment real fast.
Edit: ooh CSAM=bad. Don’t launch nuclear weapons. I could go all day.
Everything is ideological.
With nuclear weapons you could conceivably go to a small list of specific places to inspect stocks of weapons. You can watch for tests from outer space or measure the quakes they create.
My read of the “deep seek moment” was that a lab came out of nowhere and trained a very powerful model with vastly less power than was previously required.
The AI labs are private, not owned or run by the government. They are distributed widely within and across countries. The barrier to entry is much lower than that for nuclear weapons.
It’s just a complete non-starter.
(And yeah I’m deliberately ignoring the obvious incentives Dario has to recommend this course of action as the owner of a leading lab himself.)
Which is also ironically why even anti-ai people distrust any warnings of the potential of ai as an extinction level event, even if that warning would be legitimate and we ought to take seriously we are regardless impotent to do anything about it no matter what.
Well your first mistake is to believe democracy existed before AI, it just leaves you as a warrior for the status quo of the before times, which is what made tech feudalism inevitable in the first place.
What I really wanted to point out though is that for all the leaders asking for a slowdown, there are probably around 50-100 technical people who are crucial to moving the frontier forward. So, if you're one of these people, just stop. Seriously, you're already rich. Just take a vacation or get knee surgery or whatever.
Sure, I'm joking a bit, but I just write this because I see all these folks high up in OpenAI and Anthropic writing as if they have no agency. The number of folks with the technical chops to really push on the forefront of AI is just not that big. I'm not saying other folks wouldn't eventually step up, but a 6 month to a year slowdown could go a long way to reducing risk. And it's not like you need to totally stop, just say you'll only work on interpretability or whatever else is a risk reducing endeavor.
All these brilliant people who are acting like automatons between writing their scary blog posts.
> there are probably around 50-100 technical people who are crucial to moving the frontier forward. So, if you're one of these people, just stop.
I don’t think this is true. I doubt the technical know-how is nearly as much of an impediment to progress as raw compute. If AI were something that could be trained end to end on consumer hardware then crowd powered open source would blow the “labs” out of the water. The moat is money, not ability or innovation.
Personally I agree though. I would not be able to work on AI model development right now as I just don’t think it’s ethical. Fortunately for OAI/Ant I lack the relevant skillset anyways.
Are you talking about the Jacob Coxon story and the people that came out agreeing with his post? It was a well coordinated campaign, not a genuine grassroots development.
Utter bullshit. We are still scaling transformer architectures initially introduced ten years ago. Ten years of phds trying to improve on what Google produced ten years ago, mostly failing.
What has actually changed and can explain the progress we’ve seen? Do you really think it’s just a collection of better and badder RL gyms? Get real!
The main thing driving progress in ai is massive amounts of capital investment and incremental improvements in hardware (mostly memory bandwidth/capacity) and computer networking (we got better at collectives). The AI labs, ironically, have nothing to do with it. They’re the vessel for capital. The people actually driving things forward mainly work at nvidia.
The reason the models are better today than they were 3 years ago is almost exclusively due to better hardware and infrastructure software. Not better data, not better model architecture. The evidence of this is pretty easy to feel: the reason opus 5 doesn’t feel much more capable than opus 4.6 did is because they run on the same hardware generation. The reason opus 4.6 felt much more capable than anything before it is because it coincided with the scale out of a new hardware generation.
I'd like for more commments to contribute at least small things to the discussion instead of just mindless parroting.
Also, I can't follow your advice, as I don't work in AI at all ;)