GPT-6 Astra vs Claude Fable 5.1: The 72 Hours That Reset the AI Race

Anthropic shipped Fable 5.1 and Mythos 5.1, an Azure failure knocked most AI assistants offline, and OpenAI launched GPT-6 Astra. Here is what actually changed, with verified numbers.

9 min read

What actually shipped in those 72 hours

Between September 1 and September 3, 2026, the two labs at the front of the industry both shipped a flagship model, and in the middle of it most of the world briefly lost access to every assistant they use.

Anthropic went first. On September 1 it released Claude Fable 5.1 and Claude Mythos 5.1, describing them as the same model with different levels of safeguards. Fable 5.1 is generally available. Mythos 5.1 is not.

Two days later OpenAI launched GPT-6 Astra and told the world it had entered a new phase. Greg Brockman's framing was blunt: "Welcome to the AGI era."

Between those two announcements, on the morning of September 3, Azure's East US region failed and took ChatGPT, Claude, Grok, and Microsoft's own Copilot down with it.

Anthropic introducing Claude Fable 5.1 and Claude Mythos 5.1

Anthropic introducing Claude Fable 5.1 and Claude Mythos 5.1

That sequence is the story. Two labs claiming the frontier, priced identically, separated by a few benchmark categories, and both running on infrastructure that failed at the same time.

The morning every AI assistant stopped working

The failure started at roughly 6:58 a.m. ET on September 3. By 7:53 a.m. ET, Downdetector had logged more than 5,000 ChatGPT reports, and volumes peaked around 11:00 a.m. ET.

ChatGPT drew more than 37,000 reports on its own. Counted together with Codex, the figure passed 66,000. OpenAI's status page listed 15 affected ChatGPT components, including conversations, login, search, file uploads, voice mode, image generation, Deep Research, and Agent.

Claude peaked at 1,324 reports and Grok at 1,365. Both recovered inside about half an hour. Everything was back to normal by 12:42 p.m. ET.

The interesting part is who stayed up. Google Gemini logged around 500 reports and kept running, because it sits on Google's own infrastructure rather than a third-party cloud. Microsoft Copilot, running on Microsoft's own cloud, did not stay up, which points at shared infrastructure layers rather than any single customer's configuration.

For anyone shipping on top of these APIs, that is the lesson worth keeping. Three competing vendors failed together because they shared a region, so multi-model routing is only real redundancy if the models sit on different clouds. It is the same concentration risk driving the current wave of AI data center expansion.

GPT-6 Astra and the AGI claim

OpenAI positions Astra as a generational jump rather than an increment, and the pitch centres on computer use. The model navigates software the way a person does: filling forms, updating CRM records, managing calendars, operating spreadsheets and Power BI, and driving engineering tools such as KiCad and FreeCAD.

OpenAI announcing GPT-6 Astra, a new generation of intelligence

OpenAI announcing GPT-6 Astra, a new generation of intelligence

The numbers behind that claim are strong. Astra scores 72.6 percent on the OSWorld 2.0 offline subset at roughly 40 minutes per task, against GPT-5.6 Sol at 65.7 percent and about 75 minutes. That is a better result in a little over half the time. On ScreenSpot-Pro it reaches 92.7 percent.

Brockman's explanation for why this matters was that the bottleneck was never the model. "We've been bottlenecked by people writing connectors," he said. A model that can drive a normal desktop does not need one.

The safety side of the launch is louder than the capability side, though. Astra is the first model OpenAI has classified at the Critical cybersecurity threshold under its Preparedness Framework, meaning it can autonomously find unknown vulnerabilities and build exploit chains. It found two previously unknown zero-days during evaluation and scored 100 percent on ExploitBench.

Chief scientist Jakub Pachocki paired the AGI framing with a caveat worth repeating: "Progress in intelligence does not guarantee progress in alignment." That tension is exactly what AI governance work in 2026 has been trying to get ahead of, and it is now a product constraint rather than a policy question.

Fable 5.1 vs GPT-6 Astra: how the benchmarks compare

Neither model sweeps. Leadership swaps by category, which is a meaningful change from the previous generation covered in our GPT-5.6 versus Fable 5 comparison.

BenchmarkGPT-6 AstraClaude Fable 5.1
Artificial Analysis Intelligence Index6166
Coding Agent Index67 (Codex)70 (Claude Code)
FrontierMath Tier 4 v297.6%87.8%
Humanity's Last Exam (with tools)57.2%65.0%
Terminal-Bench 4.057.7%55.8%
DeepSWE v1.174.1%69.9%
OSWorld 2.072.6% (offline subset)77.9% (partial)
GPQA Diamond96.0%Not published
ExploitBench100%Not published
Context window1M tokens1M tokens

Read that table by column and Astra looks like the maths and computer use model, while Fable 5.1 looks like the reasoning and coding agent model. Astra also reports 88.0 percent on SRE-Bench in one attempt and 96.3 percent on long context in the 512K to 1M band, and it effectively saturates ARC-AGI-3, reported between 98.6 and 99.9 percent depending on the harness.

Anthropic's own standout number is elsewhere. Fable 5.1 hits 52.6 percent on Terminal-Bench-Science 0.1, against 24.7 percent for Fable 5, 29.0 percent for Opus 5, and 22.4 percent for GPT-5.6 Sol. Mythos 5.1 pushes Terminal-Bench 4.0 to 60.9 percent, ahead of both public models.

The customer signal matters too. Cognition said it was moving its Opus 5 traffic in Devin to Fable 5.1 on launch day, and Jane Street reported that Fable 5.1 solves more of its coding problems than either Fable 5 or Opus 5.

What these models actually cost

The headline rates are the same, which makes the surrounding structure the real decision.

Cost lineGPT-6 AstraClaude Fable 5.1
Input per 1M tokens$10$10
Output per 1M tokens$50$50
Fast or premium mode~$20 in / ~$100 outNot offered
Cache reads per 1M tokensSeparate rate, unpublished$0.25
Blended price per 1M (Artificial Analysis)$7.70 at max effort$7.17

Anthropic cut cache reads by 75 percent, and that single change is what moves the bill. It says typical workloads land about 25 percent cheaper than Fable 5, and heavily agentic workloads up to 45 percent cheaper. If your application replays a long, stable prefix on every turn, that is the number to model.

OpenAI argued the opposite framing. "Pricing tokens doesn't make any sense," Brockman said, pushing price per completed task as the honest metric instead. There is something to that, given Astra finishes computer use tasks in roughly half the wall clock time of its predecessor.

Both positions are self-serving, and both are partly right. Measure cost per completed task on your own workload, then check whether cache hits are doing the work you think they are.

Mythos 5.1, watermarking, and who gets the keys

The most consequential detail of the Anthropic launch is not a benchmark. It is that the strongest configuration is not for sale.

Mythos 5.1 runs the same model as Fable 5.1 with safeguards tuned for cybersecurity and life sciences, work the general safeguards would otherwise refuse. Access is limited to vetted US organisations through the Cyber Verification Program and the Life Sciences Verification Program, the latter run in partnership with the US government.

Anthropic also reports that the general safeguards got less annoying. Cybersecurity classifiers fire with 60 percent fewer false positives, and biology safeguards trigger 85 percent less often on benign requests than Fable 5 did.

Then there is provenance. Fable 5.1 embeds an invisible numerical watermark in generated text, with a detection API in private preview for regulators, law enforcement, researchers, media, fact-checkers, and educational institutions. That is EU AI Act groundwork, and it is the first time a frontier lab has shipped text watermarking as a default rather than a research demo.

On the OpenAI side, access is staged too. Astra opened through a gated programme for enterprise customers before rolling out to ChatGPT Plus, Pro, Business, and Enterprise, with a separate GPT-6 Astra Pro tier for the higher plans. Enterprise admins enable it per workspace, and it is off by default.

What to do with all this if you are shipping

The practical takeaway is smaller than the headlines suggest, and more useful.

First, stop treating the leaderboard as a purchasing decision. When two models split categories this evenly, your own evaluation set is the only thing that resolves the tie. Run both against the tasks you actually serve, at the effort levels you can afford.

Second, price the cache, not the model. Identical per-token rates mean the winner on your invoice is decided by cache read pricing and how well your prompts hold a stable prefix. That is a prompt architecture problem before it is a vendor problem.

Third, treat the September 3 outage as a design requirement. If your fallback model shares a cloud region with your primary, you do not have a fallback. The teams that stayed online were the ones already routing across providers on different infrastructure.

Fourth, plan for gated capability. Between Mythos 5.1's verification programmes and Astra's Critical cybersecurity classification, the strongest models now come with access review attached. If your roadmap depends on frontier security or life sciences capability, the paperwork is now part of the timeline, which is a shift teams following agentic coding in 2026 should budget for.

The 72 hours did not settle which lab is ahead. They settled something more useful, which is that the answer now depends entirely on what you are building.

Rune AI

Rune AI

Key Insights

  • Anthropic released Claude Fable 5.1 and Mythos 5.1 on September 1, then OpenAI launched GPT-6 Astra on September 3, putting two frontier releases inside 72 hours
  • An Azure East US failure on September 3 took ChatGPT, Claude, Grok, and Copilot offline for roughly six hours end to end, while Gemini stayed up on Google's own infrastructure
  • Both flagship models list at 10 dollars per million input tokens and 50 dollars per million output tokens, so the real cost difference comes from Fable 5.1's 0.25 dollar cache reads and Astra's 2x fast mode
  • Benchmark leadership is split rather than decided: Astra takes FrontierMath Tier 4 and computer use, Fable 5.1 takes Humanity's Last Exam with tools and the Artificial Analysis Intelligence Index at 66 to 61
  • GPT-6 Astra is the first model OpenAI has classified at the Critical cybersecurity threshold under its Preparedness Framework, and it found two previously unknown zero-days during evaluation
RunePowered by Rune AI

Frequently Asked Questions

Is GPT-6 Astra better than Claude Fable 5.1?

It depends on the task, and the honest answer at launch is that neither wins outright. Astra leads on maths and computer use, scoring 97.6 percent on FrontierMath Tier 4 v2 against Fable 5.1 at 87.8 percent, and 72.6 percent on the OSWorld 2.0 offline subset. Fable 5.1 leads on reasoning breadth and coding agents, taking Humanity's Last Exam with tools at 65.0 percent against Astra at 57.2 percent, and scoring 66 to Astra's 61 on the Artificial Analysis Intelligence Index. Terminal-Bench 4.0 is close enough to call a draw at 57.7 percent for Astra and 55.8 percent for Fable 5.1.

What caused the September 3 AI outage?

A failure in Microsoft Azure's East US region, which took down ChatGPT, Claude, Grok, and Microsoft's own Copilot at roughly the same time. The trouble began around 6:58 a.m. ET and services were back to normal by 12:42 p.m. ET. ChatGPT alone drew more than 37,000 Downdetector reports, rising past 66,000 once Codex was counted with it. Google Gemini stayed largely up because it runs on Google's own infrastructure rather than a third-party cloud.

How much does GPT-6 Astra cost compared to Claude Fable 5.1?

The headline rates are identical: 10 dollars per million input tokens and 50 dollars per million output tokens for both models. The difference is in the extras. Astra offers a fast mode at roughly double the price, about 20 dollars input and 100 dollars output. Fable 5.1 cut cache reads to 0.25 dollars per million tokens, a 75 percent reduction that Anthropic says lowers typical workload costs by about 25 percent and agentic workload costs by up to 45 percent. On blended pricing, Artificial Analysis puts Fable 5.1 at 7.17 dollars per million tokens against Astra at max effort at 7.70 dollars.

What is Claude Mythos 5.1 and can I use it?

Mythos 5.1 is the same underlying model as Fable 5.1 running a different set of safeguards, tuned for cybersecurity and life sciences work that the general safeguards would otherwise block. Most developers cannot use it. Access is limited to vetted US organisations through two programmes, the Cyber Verification Program for defensive security work and the Life Sciences Verification Program for research professionals. Fable 5.1 is the generally available model and is what you get through the API, Claude.ai, AWS, Google Cloud, and Azure.

Conclusion

Three days is not much time to absorb two frontier launches and a multi-provider outage, but the pattern underneath them is clear enough. The capability gap between the top two labs has narrowed to the point where benchmark leadership swaps by category rather than by model, which means the choice in front of most teams is no longer about which lab is ahead. It is about which model is cheaper for the specific shape of work you run, and whether you can still serve traffic when a single cloud region has a bad morning. Astra's Critical cybersecurity classification and Fable 5.1's watermarking both point the same way too, toward a period where access control and provenance matter as much as the scores on the chart. Pick on your own evaluations, price the cache reads honestly, and put a fallback behind whichever model you choose.