GPT-6 Astra is OpenAI’s frontier model, released on September 3, 2026. It operates software on your behalf. It browses, fills out forms, updates CRM records, works inside document editors, writes and tests code, builds and QAs websites, and runs multi-step workflows while you do something else.
For most business owners, the useful move is to set your systems up so you can switch models quickly. The top two models are already trading places depending on the task, and that will keep happening.
At TJ Digital we run AI through every part of our process for 40+ clients. We can consider moving most of our work to GPT-6 within a week of launch because our clients’ brand knowledge lives in documents we own, which means a new model is a decision instead of a rebuild.
I spent the last day reviewing the benchmarks, the demos, and the reactions from people who had early access, and I’m shocked. I understand this sounds like hype to anyone who hasn’t been following closely. I still think GPT-6 marks a new era in AI, and I don’t think it’s an overstatement to call it the AGI era.
Table of Contents
ToggleWhat Is GPT-6 Astra?
GPT-6 Astra is OpenAI’s agentic computer-use model, rolling out to paid ChatGPT tiers, the OpenAI API, Microsoft Azure and AWS Bedrock. OpenAI describes it as being able to operate browsers and desktop interfaces directly.
The shift here is from answers to action. Plenty of AI deployments over the past two years still left the execution to a human. The model researched the prospect and a salesperson pasted the results into Salesforce. The model wrote the article and an editor published it. Astra moves a lot of that work from recommendation to execution, inside whatever permissions you set.
@tjrobertson52 What GPT-6 Astra means for your business: Fable got 30-something on ARC-AGI-3 last month. GPT-6 got 99%. #GPT6 #AI #SmallBusiness #AGI
♬ original sound – TJ Robertson – TJ Robertson
What Can GPT-6 Astra Do on a Computer?
Computer use is the thing AI has been worst at, and it’s where Astra posts its biggest gains.
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Best Claude result |
| OSWorld 2.0 (offline partial) | 72.6% | 65.7% | 70.2% (Opus 5) |
| ScreenSpot-Pro, no tools | 92.7% | 76.9% | 87.3% (Fable 5) |
| AutomationBench | 41.4% | 18.1% | 31.4% (Fable 5.1) |
Speed moved just as much as accuracy. Astra completes a simulated OSWorld task in roughly 40 minutes versus about 75 minutes for GPT-5.6 Sol, a 47% reduction. On Mind2Web, OpenAI reports 1.9x faster task completion.
That speed number may matter more commercially than the capability number. An agent that succeeds but needs an hour of wall clock time and three retries has terrible economics. Cutting a task from 75 minutes to 40 changes what you can run while an employee waits instead of scheduling it overnight.
Keep that AutomationBench score in mind before you hand an agent your ad account. 41.4% is more than double the previous generation, and it still means the model fails most of the time on a benchmark built specifically around automating professional workflows.
Did It Really Score 99% on ARC-AGI-3?
Yes, with an asterisk that matters. OpenAI reports 99.9% on ARC-AGI-3. ARC Prize’s own testing records roughly that same figure using OpenAI’s provider-adapted agent setup, and 62.7% using ARC Prize’s standard test configuration.
Both numbers are real. The score reflects the whole deployed system, including the scaffolding wrapped around the model. Context management, tool interfaces and execution setup change measured capability on their own, which is why the model plus its agent layer is becoming the actual product.
The more interesting finding came from ARC Prize’s Greg Kamradt, who said Astra surpassed their human action-efficiency baseline on 96% of levels. Astra solved most levels using about as many actions as a person would.
One correction to something I said in my video: the roughly 30% ARC-AGI-3 score from a month earlier belongs to Claude Opus 5. Claude Fable 5.1 has no published ARC-AGI-3 result. The jump is still enormous. The attribution was mine and it was wrong.
Is Astra Better Than Claude Fable 5.1?
Not across the board, which is what I assumed at first. Artificial Analysis found the two tied at 53 on its Intelligence Index and tied at 62 on its Coding Agent Index. Astra hit that Intelligence Index score at about 40% of Fable’s measured cost per task ($3.26 versus $7.63), using roughly 27,000 output tokens per task versus about 78,000.
Fable wins outright in places too. On Humanity’s Last Exam with tools, Fable 5.1 scores 65.0% against Astra’s 57.2% in OpenAI’s own comparison table.
Here’s how I’d read the evidence by type of work:
| Type of work | Evidence currently favors | Why |
| Operating a browser or GUI | GPT-6 Astra | Strong OSWorld and ScreenSpot results plus large speed gains |
| SaaS and workflow automation | GPT-6 Astra | 41.4% vs 31.4% for Fable 5.1 on AutomationBench |
| Terminal and coding agents | Very close, slight Astra edge | 57.9% vs 55.8% on Terminal-Bench 4.0, tied on independent testing |
| Advanced math and scientific terminal work | GPT-6 Astra | Large leads on FrontierMath Tier 4 and Terminal-Bench Science |
| Tool-assisted reasoning | Claude Fable 5.1 | 65.0% vs 57.2% on Humanity’s Last Exam |
| Prose and brand writing | Claude Fable 5.1 | Canva, Glean and other production users report stronger writing quality |
| Slides and design taste | Unresolved | Astra’s design benchmark doesn’t include a Fable 5.1 result |
| Cost per completed task | GPT-6 Astra | Same aggregate scores at materially lower measured cost |
I said in my video that Fable is probably still better at writing and may have better taste in design. The honest version is that nobody has run a clean public head-to-head on either. Anthropic’s launch materials include customer evaluations praising Fable’s prose and its decks, and Canva’s head of AI called writing the standout in Fable 5.1. Those were blind comparisons against Fable 5, not against Astra.
Why Is Safety Slowing OpenAI Down?
This is the part that made me feel like we’re in a new era, more than any benchmark did. Our top concern is shifting from how capable the models are to how safe they are.
OpenAI President Greg Brockman said in a September 4 interview that safety, security and alignment are becoming “almost the bottleneck” to development. He singled out monitorability as the key requirement as models get more capable.
The operational record backs it up. OpenAI says it paused frontier training, including Astra training, for two weeks after a security incident. It then held back larger reinforcement learning runs and only restarted a major one on August 28, 2026 once higher safety thresholds were met. In the week after preliminary evidence suggested Astra might hit OpenAI’s Critical cybersecurity threshold, Astra-class GPU allocation dropped another 59.2% while allocation to other model classes rose 17.2%.
The cyber capability explains the caution. OpenAI classifies Astra as its first broadly deployed model to reach the Critical cybersecurity threshold, and reports that it discovered two previously unknown zero-day vulnerabilities during evaluation.
There’s a real paradox in the safety data. Astra misbehaved on 2.4% of OpenAI’s computer-use safety tests versus 22.0% for Sol, so it appears better aligned. OpenAI also says Astra’s written reasoning is harder to monitor, partly because the model controls how much of its reasoning it exposes on simple tasks.
I said OpenAI could ship a better model every week and safety is all that’s holding them back. I’ve since looked for a first-party source on the weekly cadence and couldn’t find one, so treat that part as commentary. The safety bottleneck itself is well documented.
What Should Businesses Do Now?
My advice is the same as it’s always been. Make sure your knowledge base and systems are set up so it’s easy to switch from one model to another. Astra and Fable 5.1 launched two days apart and already trade places depending on the task, and single leaderboard snapshots go stale within days.
What portability actually requires:
- Keep your knowledge in documents you own. Customer facts, brand rules, product data, SOPs and campaign knowledge belong in model-independent files you control.
- Put an adapter between your business logic and the model. Avoid hardcoding a model name throughout an application.
- Build task-specific evals. Keep a regression set of real jobs (briefs, landing pages, decks, CRM updates, report QA) and re-run it every time a model changes. Generic benchmarks don’t measure your brand.
- Route instead of standardizing. Send computer-operation work to whichever model wins on your own tests, and keep the other one where your reviewers prefer its writing.
- Measure cost per successful task. Track retries, latency, human review and failure rate per completed outcome. Dollars per million tokens tells you very little on its own.
- Use least-privilege permissions. Separate read, draft and publish access. An agent that can take real actions can take real wrong actions.
- Keep a human on anything irreversible. Ad spend, production deploys, client emails, external publication, account changes.
- Log everything the agent does. Tool calls, changed records, model version, approvals, outcomes.
For marketing specifically, the workflows that now look automatable are the browser-heavy production ones: research to brief, CRM hygiene, reporting dashboards, spreadsheet analysis, CMS staging, page QA. The best candidates are the ones where you can objectively check the result before anything irreversible happens.
Writing works differently. Brand voice doesn’t show up in a generic benchmark. Before you move content production to whichever model tops the leaderboard, run 50 to 100 real briefs through both and have a senior editor score them blind. That will tell you more than ARC-AGI ever will.
How We’re Using GPT-6 Astra at TJ Digital
My assumption right now is that we’ll switch most of our work over to GPT-6. We’ll keep testing Claude on writing before we move anything that touches client content.
That switch is a decision rather than a rebuild because of the Brand Ambassador we build for every client during the Two-Week Strategic Assessment. It’s a set of documents covering the client’s brand, voice, audience, services and competitive landscape, and it belongs to the client. When a better model shows up, we point it at the same documents.
That’s the part I’d push every business owner on. Everyone gets access to the same frontier models within days of launch. The durable advantage is clean proprietary context, well-designed tools, fast evaluation, and the ability to replace the model without rebuilding your process around it.
Is GPT-6 Astra AGI?
There’s never going to be a clear bright line showing when we hit AGI. What I can say is that Astra performs computer work at roughly human action efficiency on most of the levels ARC Prize tested, and OpenAI reports it completes many everyday tasks faster than the user could. That’s close enough to change how I’m planning the next year.
How Much Does the Astra API Cost?
$10 per million input tokens and $50 per million output tokens, identical to Claude Fable 5.1’s list pricing. The difference shows up in cost per completed task, where independent testing measured Astra at roughly 40% of Fable’s cost for the same aggregate score.
Should Small Businesses Switch to GPT-6 Astra?
If your AI work is browser or software operation, test it now. If your AI work is writing, test it before you switch anything.
Whatever you’re using today, the more valuable project is making your knowledge portable so the next launch costs you an afternoon. If you want help working out what your business should actually be doing about AI search, get in touch. I’ll tell you straight whether we can help.