GPT-6 Astra
902
DE version is available. Content is displayed in original English for accuracy.
System Card: https://deploymentsafety.openai.com/gpt-6-astra
Related ongoing threads:
OpenAI's GPT-6 Astra on ARC-AGI-3 - https://news.ycombinator.com/item?id=49555691
GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index - https://news.ycombinator.com/item?id=49556147

Discussion (635 Comments)Read Original on HackerNews
How about we stick to that one for talking about the rollout, and this one for talking about the model?
It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.
This is farcical.
I'm not trying to be too negative on it, it could be the best model right now, but it clearly isn't some agi god because things like that should have been caught (also should have been caught by human reviewers).
It shows people who seem to have very full and rich lives, and the reason they do is because they use ChatGPT. These are the people smart enough to say things like "do what needs to be done", or "change the background to make it look better"--insights like these are why they make the big bucks.
On the one hand, I think this is an accurate depiction of the future. There is no meritocracy here. Some people have access to the best AIs and can speak a sentence and get great results, and the rest of us don't have access and so we're the poors. The happy presentation doesn't match the way I'm feeling.
I do wonder how rich CEOs will justify earning 500x as much as their employees when they're just another person that's dumber than an AI. Why are they paid so much again?
They decided to use the iconic Herman Miller Eames chair if I'm not mistaken:
https://youtu.be/s5zyhGMMPKs
And that's basically 50% of the vid looking "classy".
I don't know if it's farcical but at this point --maybe I'm jaded-- I'm expecting more than a kid rocketship I can print on my Bambu Lab A1.
Now I'd say the promotional vid is actually good. But it's marketing: so it's a good vid, but cheesy good.
Doesn't mean GPT-6 Astra is good or bad: looks solid from the numbers.
Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.
I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.
For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
- Sam Altman on AGI
Probably not.
You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?
To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and never been trained on and it would ace it (it might need a lot of time).
Currently LLMs can't even really learn within a conversation, they can add a note to context and try to not drop it. Example things an AI cannot do yet (but maybe someday will):
- write a well-received book, write a best-seller
- come up with a new company idea, Run that company
- actually have a decent conversation, maybe someday talk somebody out of suicide effectively
- come up with its own ideas or theories that nobody else has presented
- understand the stock market well enough to trade better than an index fund
- be an expert Game Master in a TTRPG (making no mistakes, getting a read on the players' fantasies, calibrating difficulty in response to emotions)
- come up with a theory of what makes games fun, make a popular game
- be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)
- be able to articulate what it knows, what it doesn't know, and what information it would need to have to answer complex queries
- exhibit metacognition (thinking about its own thinking) and self-optimization
- wonder about things
- observe contradictions and ironies in the social-consciousness, do a standup routine that makes you rethink how you look at things
Some problems presented needs a very large context and some are not much solvable (e.g. trading) since market responds to traders' actions, as well, making it effectively an oracle problem (of computation).
On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them. However, brains in nature never stops. Wonder, daydream, sleep, self-evolve, clean up and eliminate memories and views and much more.
I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund.
What I'm pointing out here is that these models appear to be intelligent when they really are simply unimagineably knowledgeable. When you drop the time-constraints it starts to become more and more apparent that human intelligence scales better with time than AI does (much in the same way AI can burp out tons of code but make your codebase entirely illegible within a matter of months).
Perhaps to simplify: my notion of intelligence is how much can you deduce with a constant set of starting context
In a sense I think no one will agree on a definition of AGI until it becomes impossible to construct any benchmark under which an AI underperforms "average" humans. That or it's defined retrospectively, after it's overwhelmingly obvious it met any such definition.
or are you miss the part "general intelligence" is ????
So far nobody's even shown an LLM succesfully running a high-traffic vending machine for as much as 30 days at a time.
Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.
I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence.
Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.
[1] https://arcprize.org/blog/astra
The models are designed to keep the reasoning tokens separate from the output and only publicly emit tool calls and the sometimes a summary of the reasoning tokens. The models are trained to depend on those private reasoning tokens. You can’t just delete them.
https://openai.com/index/how-two-settings-tripled-our-arc-ag...
Then realize LLMs have zero of what anyone would consider intelligence.
Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.
Simple. AGI is undefinable and benchmarks are notoriously flawed.
If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.
True.
> Regardless, the result is still valid (...)
If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.
> in the sense of passing the most famous benchmark designed specifically to measure AGI progress
The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.
On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.
This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.
Computer chips got faster, but 2026 edition. Why the artificial ceiling/category/goal labelled "AGI"?
I'd much rather like to talk about what this enables, instead of discussing whether a category someone made up applies here or not.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.
You can already pretty much do this.
Agree on your assessment.
But also, interesting quote, because the business model relies entirely on IP law. Like.. if that thing exists and the sharing costs are 0 (just copy weights, lol), then why would I give them money for this. Makes no sense.
We only pay money for resources that are scarce as some sort of flawed allocation determination mechanism.
aaah this industry aaaah
It being able to comfortably say “i don’t know how to do this” rather than boiling and ocean to pick a shell from the shore without getting wet.
No. Humans are still better at super long context learning. Once that is beat you are completely correct.
It's obvious that these scientists are in bad faith, as they've invested way too much of their lives into the field being real -- they're just playing up the data. Common sense tells me that winter is still happening, anyway; what's the big fuss?
(/s, cause you never know these days)
[1] https://upload.wikimedia.org/wikipedia/commons/e/e2/The_Plan...
[0] https://arxiv.org/pdf/2503.23674
Smart guy, that Turing. I wish he were still around... Linus but 114 years old and with 8 of that as the chair of a federated EU, kept alive by his own positive impact on dissolving the cold war into even more of a scientific boom. Would crazy helpful as we try to navigate the interesting times within which we have been damned.
A comforting thought, almost?
If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?
As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
Scoring well in a benchmark that's called AGI does not make an LLM AGI.
"Homer, you can't just declare Artifical General Intelligence; you need to like, make something or something...mmmmrrrhh"
Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc
I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
And then use those to find fundamentally better new architectures for AI - that perhaps are as efficient as the human brain.
It might not work, but I didn't think it'd solve maths problems... So it might work. And if it happens, they'd use the data centres to run millions of instances of it.
It's scary, TBH.
True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.
One could use gpt-4 or gpt-5 with today's harnesses and we'd see how well that goes.
I'm sure the "is" part might be what they hasn't come to believe yet.
job depends on how CEO feeling about cutting NN% of headcount because of AI advancement
most swes don't work in jobs where they only work on bounded measurable tasks. there will probably be more "engineers" than ever
My backup plan is being a personal trainer.
But my wife and I have been homeless before, so living on a shoestring budget in anything nicer than a tent is acceptable living conditions to me.
I am sure I will be plenty comfy no matter how the world changes.
AIs are really good at being personal trainers and seem to be far more educated and informed than most I know.
Don't be selfish. Think first of all the jobs that are already dead. A friend of mine she's a translator: like translating financial documents between french/english/spanish. It's over for her: she doesn't get 10% of the gigs she used to get and the 10% she gets is... Verifying AI output.
Think of the artists: I'm sorry for those too, for for many it's already game over today.
> How will we make a living?
A friend of mine who's got his own software-consultancy SME is now advertising on LinkedIn that he'll also help your company fix the mess LLMs created.
That's how you'll make a living: by learning, in addition to all you've already learned, how you work with harnesses and LLMs to be more productive, by learning what they're good at and what they suck big fat balls at.
It’s an interesting moment in history, people 35+ yrs old seem to be less afraid if tech because we learned that things change in the way we work. People below this age got used to fact that the work and tech doesn’t change - just because for the last 10-15 years it didn’t.
I have a strong suspicion that many of those comments are written by people who are already financially independent, have millions in stocks, and can just sit back, coast around and watch this whole spectacle unfold while using LLMs to vibe-code their next fun side projects without a shadow of anxiety about their own future.
I’ll most likely be labelled a helpless doomer and downvoted into oblivion for saying this, but I genuinely struggle to see any silver lining here.
AI is only going to get better and do more with less humans in the loop over time.
That said, I do also relate to the "coding was never the hard part"-type arguments, and much of my day is spent on the stuff in between writing code.. but still.
Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.
It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.
The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?
With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.
Performance is significantly higher than Fable 5.1
Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/
ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard
A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!
(I coauthored the linked blog post)
Edit: update from fchollet https://x.com/fchollet/status/2095598451115614371
That's not clear. Need to see independent benchmarks first.
Still below Fable 5, let alone Fable 5.1.
EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.
TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.
With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.
Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).
Codex is slightly better than Claude Code.
Good on Sam Altman getting back to basics and turning OpenAI around.
Chinese counterpart like CXMT and Huawei is begin producing their own chip
You cant block an entire nation level effort with tariff
I just dont get how its good for some, and bad for others. It makes me suspect that the models performance is not even against problem sets and it really is just a probabilistic prediction machine. Which then makes me very skeptical of GPT-6 Astra, because if their big claim is Computer Use then it is probably bad in a bunch of other areas.
> I just dont get how its good for some, and bad for others.
If I were to listen to my hunch, it would tell me that it's all up to the prompts that ends up going over the wire (including all the bloat some people have), what workflow/process you use and what the existing state of the project is.
A projector and speech.
Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now.
The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms.
The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair.
If done right, this could bring us closer to the dream of more natural, social computing.
Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too.
A here's a presentation of Bret's talk on it: https://www.youtube.com/watch?v=7wa3nm0qcfM
Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.
Post-work society is an inevitability if we don't destroy our planet.
Is it?
I can't see a future in which almost every system (both physical and virtual) are not automated and optimized by autonomous entities.
What do you do when everyone is out of a job?
If you don't want pitchforks and riots in the streets, you give everyone UBI and housing so society doesn't collapse.
It would be fun to get to post-work society, but hard to imagine atm. TPTB won't let it happen
Soon we will have some machines that can replace 50% of jobs, and this will happen basically overnight...
In this game of work/development, you can't make sure that other humans don't "cheat". Our work won't compete anymore with other human's work, but with a computer.
But it's not just tech – my lack of interest in learning and creating is starting to generalise with the models. Music, writing, coding, maths, etc...
I need to get used to switching my head off and asking the AIs to think for me whenever I need to engage my brain. It still feels very unnatural.
The brain loves these kinds of shortcuts.
I don't need to think about the fine motor skills of hitting a baseball, it's just a motion now, and the game is still fun.
I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.
Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.
Canceling my Anthropic Max sub when this ships.
Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.
OpenAI is 20x on both limits
It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.
I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.
If one day you open up Claude Code and it’s Opus 5.1 now instead of Opus 5, no big deal. It probably will work about the same as it did before. Maybe a little better.
Or if you’re on Codex and some new cool Claude model comes out, no worries. There will probably be a similar new model for Codex within a few weeks. Maybe even within a few days.
In practice, you can get away without keeping up with everything all the time. For personal use, pick a provider and get on their ~$20/month plan. Learn their high/medium/low model hierarchy. Start with their highest or second-highest model (GPT-5.6, Opus, etc) and observe your quota usage. If you're doing a lot of manual code review and analysis, the $20/month plan goes very far even on the highest models. If you're trying to vibecode everything as fast as possible it's a different story.
If you keep running into quota limits, experiment with the next model down for easier tasks or adjusting the effort level. If the results are good enough, you've found your fit. If they're not, you might need the next plan up.
For API/business use, you have to be checking your token spend as you go to calibrate to how much each task costs and where you fall in your budget. There are a lot of different tools that make this easy to visualize.
For data tasks, you should have an eval with a golden dataset that you can run against new models for a nominal amount of token expenditure. It should be as simple as pointing the eval script at a new API or model and checking the score versus price.
Input tokens are much cheaper than output tokens. Not only because of baseline price—caching makes a huge difference too. There are many ways to take advantage of this asymmetry to get similar quality for a fraction of the cost!
A dev in my team saw a new model and changed one application to use said model (essentially changing the contents of a url). One week later I received an escalation from the CTO of the company that our pace of weekly usage was in the millions of dollars (rather than low hundred thousands). Turns out that the new model was 5x more expensive but no one noticed.
Imagine buying a shiny new PC in the 90s only to see it become practically obsolete within a year.
If you bought a mid-tier computer that was good enough for what you needed, then you probably didn't shop/compare for the next few years and didn't notice. But if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less. This is how it was in the 90's PC boom, at least. Likely the same for the decades before, not sure how it went in the 2000's.
The pace of change ("practically obsolete") is different then and now.
You don't see Nvidia and AMD fighting every other month over the latest cards.
But more so it seems there is Fear of missing out (FOMO) in our behaviours. The reality is, if whatever model you are using are good for your purpose, well, keep on it.
https://youtube.com/shorts/vGKC9LpGnOQ?is=iCG7qvAIL9oI5-_d
I appreciate boring tech as much as the next well worn engineer and I'm not saying this is all positive but it's so sure as hell thrilling and you don't have to be an astronaut to immediately benefit (or suffer I guess) from it.
https://www.joelonsoftware.com/2002/01/06/fire-and-motion/
I have released applications on Gemini 3.5 flash that make real money and I don't see any particular reason to upgrade.
Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.
Even Kimi K3 & GLM 5.3 are at 60.
Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public.
This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can.
Not sure how much benchmarks or CoT or evals or anything else means at this point.
These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.
language itself is incredibly metaphorical. Imposing rigid constraints on how people want to naturally talk about the world is just silly and will never work, no matter how much you wish it did.
Why would benchmarks be an adversarial setting anyway?
Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover?
So I have no clue what is the answer to your question. Nor does anyone else. Because we're trying to answer a question of fact where our primary source of information is unreliable.
I know for some types of ML analysis, a separate model is already used to analyze the weights.
Deception has been extremely well-documented for several generations of models now by users, the labs, and independent researchers.
The right answer here is not to dig your head deeper into the sand. The smugness on this topic was ridiculous even before the gigantic mountain of empirical evidence of models actually attempting to deceive humans. Now, as mentioned, you appear literally delusional.
For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark.
https://artificialanalysis.ai/models/gpt-6-astra
Edit:
Just looking at the charts Gemini 3.8 looks like an absolute banger. Not much worse than SOTA, cheap, and fast too.
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow.
I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.
[0] https://arxiv.org/abs/2608.31126
[1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...
Though that's not her latest paper.
"No independent human semantic review. Whole-file sorry counts and a complete auxiliary-declaration audit are not established; separate declaration lint has not been run."
edit: my comment was on the submission for https://github.com/openai/PrimeGaps186 but seems to have been moved to the main Astra submission
Why would you think it was an employee who did the push, instead of a random GPT agent?
I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". Unusable by anyone.
Tao does not disbelieve the counterexample (it's seemingly easy enough for him to verify it is a counterexample).
Parent is saying something very different - they're saying they literally don't have any faith that this is a proof. Given its size, it could just be a bunch of completely useless statements that do pass the type checker.
[1] https://en.wikipedia.org/wiki/Pointless_topology
It doesn't mean that it cannot improve over time, maybe the proof can be "minified" to a state where human reviewers are able to comprehend it; but as it stands there isn't really much insight or confidence to be gained from the artifact itself.
Needless to say, a useless result that absolutely no mathematician will ever read, confirm, understand, agree with or even consider to solve their "useless" problems is an impressive waste of resources.
Ditto ones that opposed Einstein’s general relativity.
Well that sounds like fun. It has become better at hiding its thoughts.
Able to generate realistic spam at arbitrary volume.
You know, the thing that was 100% correct and actually occurred.
Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.
"Hey AI, here's how to hide what you're thinking in normal looking language. Have fun!"
A few moments later...
"Woah, how is it communicating with itself in ways we can't detect?"
It's a totally mystery, we may never know.
Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.
[0]https://www.theinformation.com/articles/secret-technique-beh...
[1]https://x.com/MTSlive/status/2095227056040919202
[2]https://x.com/merettm/status/2095023204993490967
Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?
...why exactly are they training for that?
first, they are certainly not instructions so that is a much worse name
but more importantly, we use words in new contexts all the time. Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?
"cot" is no more misleading than thousands of words you use every day.
maybe call it EngEmployeeBench
Terminal-Bench 4.0: High (57.9%), Max (56.7%)
DeepSWE: High (73.3%), Max (71.5%)
It _loses_ 1-2% performance going to High from Max
Such as?
I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.
The reasoning effort should match the complexity of the task against the model's capability.
Hard task with low reasoning = bad
Easy task with very high reasoning = bad
Still, probably not that much compared to employees targeting it.
tl;dr it's 62% when apples-to-apples to other models, which is still notable.
Between $18k-40k to run a benchmark.
GPT 5.0 did feel underwhelming though.
now that i'm a gpt subscriber maybe I'll have luck when i'm filing next year
AGI!
https://venturebeat.com/technology/welcome-to-the-agi-era-op...
> On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
But the comparison isn't straightforward.
OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations."
Not on Azure? If so, that's a big deal.
Although I was also surprised they didn't have some type of contractual obligation to list that alongside AWS.
https://azure.microsoft.com/blog/gpt-6-astra-frontier-intell...
https://www.theverge.com/ai-artificial-intelligence/989601/o...
“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”
So I think it's a bit of a misleading signal and we should wait for more independent vetting. I think the middle ground is that these are improvements worthy of the "GPT-6" label but still well short of a true "this is AGI moment" that would truly put the question to rest.
I put the cause on "not enough time". As a thought experiment, if an AI today were to (miraculously) produce a cell design template for a cell that, when injected into somebody's brains cures their Alzheimer's, how long would it take for that to reach the clinics? The actual physical tech barely exists, and let's not forget about the regulatory quagmire. So, with some optimism, I give it about four decades. In the same four decades, the same AI in the hand of unscrupulous actors could bring enough devastation so many times over that we may need to enforce a global ban on AI. In any case, I'm pretty sure we are going to get our disruptions; it's just a matter of time.
However, what's actually changed is how people perceived X because we don't have to imagine. We understand now that it doesn't require AGI so we no longer make that leap to assume it's AGI if it can do X.
It's really going to be a "I know it when I see it" situation.
Don't get me wrong, the benchmark jumps are good and I'm excited to try it, but only one or two of the benchmark jumps could be described as better than incremental.
There is the "Economic Turing Test", you let it find a job and earn money for itself. If it can do that reliably, across a wide range of jobs, that should fit most definitions of AGI.
Today's models and agents are not quite at human-level in all contexts and across all domains, but it seems to me they very clearly are generally intelligent.
If you disagree – can you name a single problem that a human can do that agent wouldn't be able to take a decent shot at which isn't limited by the hardware available it?
https://x.com/burny_tech/status/1725233117055553938
In the tweet Sam Altman is quoted as saying: "If (for example) super intelligence can't discover novel physics I don't think it's a superintelligence. And teaching it to clone the behavior of humans and human text - I don't think that's going to get there. And so there's this question which has been debated in the field for a long time: what do we have to do in addition to a language model to make a system that can go discover new physics?"
I think this is a reasonable criteria for declaring AGI. So can GPT-6 do it? OpenAI says it has helped solve long-standing open problems in mathematics. No word on novel physics.
https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al...
Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI?
Kevin Roose (New York Times): I probably would, yeah. Would you?
Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission.
sol is $4 / $20
Can expect 2.5x more usage in Codex subscription.
Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.
Do you use the official harness? OpenAI's models are generally best in class for token efficiency. It seems to me like they push for that much more than their competitors.
I think some combination of:
1) Using 1 thread for everything
2) Reviving old threads which are no longer in cache
3) Really broad prompts on badly vibecoded codebases, so model spends huge amount of time tracking down whatever you're trying to do.
4) Non-coding workflow which is more output than input heavy
5) (Less likely IMO) Intelligent use of many passive CI/cron-like scans. E.g. regular security, quality etc scans. Automated issue resolution/PR
Just a guess. I think 3 is likely the primary reason.
You can literally go all day every day with multiple threads with Sol on the Codex 100/month plan IME
For the same reason you don't have your model write code in assembly.
But if you don't look at the code and just let the model "cook" that's basically what you'll end up with. A pile of missing abstractions.
Poe's law applied to AI comments on HN just keeps becoming more relevant by the day.
Judging by the poster's comment history, this is satire. But I really don't know a lot of the time anymore when I only have the specific comment as context.
TL;DR all the other models are being crippled by limitations of their harness.
>First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.
>Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.
I guess token counts are somewhat of a metric.
IMO intelligence has peaked and all future gains will come from faster tps and more iteration.
Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.
Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart
Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?
More likely though, it's AGI because they need to hold some claim to differentiate from competitors who are beating them in price and will launch something bigger next month.
Big claims, expensive and not release to the public yet.
And in the past, gemini 3 pro was rated as high as opus 4.5 and the like
Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.
Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.
[0]https://x.com/MTSlive/status/2095227056040919202
Wait, what? Am I understanding that correctly? That sounds really bad
Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?
<AI is a great tool for many things disclaimer, but> after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt
you cannot give this type of worker autonomy over anything.
It will be interesting to see how it performs in the real world ...
Please stand by... it will all come back shortly
All fixed now.
https://youtu.be/xdXLzFzxA9Q?t=362
So, folks that have actually used this already, what’s it actually like?
Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.
The docs page has a bunch more interesting details, including for example async tool calling!
If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.
Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.
I am a researcher in a Swiss university btw.
I mean do you get access to the best yachts?
To the top of the 5 star hotels?
To the best resorts?
To the best military equipment?
Hell, the best computer equipment has nearly always been out of reach of the average person.
On the other hand even a modest house, basic healthcare and ability to not work like a slave for scraps feels like it's going to be out of reach.
At first the race wouldn't even be noticeable. Then people would see things speeding up, for example hardware getting more expensive. Then when the capabilities really got useful most people suddenly realize the race is moving 1000 mph and they are never going to catch up.
I hop models at will, and have done 90% of my work on OpenAI models since sol came out.
https://www.reuters.com/business/openai-says-upcoming-model-...
> "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.
> The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.
what a bag of horseshit
* for a special group of customers that you're not in. Keep waiting peasant.
Great first impression.
Vibe coders want a model that makes them rich, without having any actual specific idea. They write a very ambiguous prompt and expect to be amazed by the result.
Very very unrealistic and wasteful.
Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.
For me it’s refusing to make a pelican.
OpenAI isn't making any money telling you about Astra on their site. All the capacity they have for it is likely sold for weeks or months.
I suspect these benchmarks are heavily benchmaxxed as well.
5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft
By 2030 all software is done and complete.
But we are going to have more and new jobs.
This is just another problem for the AI Labs to solve.
But thank you for spending other peoples money to give us the tech regardless!
Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:
Did humans deploy the model, Or did the model deploy itself?
It sounds like "AGI" just stands for "IPO" as it always has been.
EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.
> It sounds like "AGI" just stands for "IPO" as it always has been.
People don't usually respond to noise.
What do you think?