Claude Opus 5.5 Beats GPT-6 Astra at a Fifth of the Cost

Anthropic's new Claude Opus 5.5 matches Fable 5.1, tops GPT-6 Astra on coding benchmarks, and cuts typical workload costs 40% versus Opus 5.

Read as article

Claude Opus 5.5 Beats GPT-6 Astra at a Fifth of the Cost

By @sharedot · · 8 pages

Anthropic's new Claude Opus 5.5 matches Fable 5.1, tops GPT-6 Astra on coding benchmarks, and cuts typical workload costs 40% versus Opus 5.

What happened

Anthropic released Claude Opus 5.5 on September 22, the first model in its new Claude 5.5 family. The company says it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5 on typical workloads, with output generation more than 30% faster. The model is available under the identifier claude-opus-5-5 on Anthropic's platforms and on AWS, Google Cloud, and Microsoft Azure. Anthropic is also raising five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans, and adding a savable rate-limit reset for subscribers.

Why it's surprising

The upset here is cost-adjusted performance, not just raw scores. At its default (medium) effort setting, Opus 5.5 scores 54.6% on FrontierCode v1.1, beating GPT-6 Astra's top score of 53.3% at roughly a fifth of the cost per task, according to Anthropic's announcement. On CursorBench 4.0 at medium effort it scores 52.5%, topping GPT-5.6 Sol's best of 41.7% by 11 points for about a third of the cost per task. Anthropic also cautions that benchmark margins have become a less reliable guide at these capability levels, noting the real-world gap to Fable 5.1 is narrower than the scores suggest — an unusual admission alongside a claim of leadership.

The evidence

Anthropic's benchmark table reports Opus 5.5 at 66.4% on Terminal-Bench 4.0 agentic coding, versus 55.8% for Fable 5.1, 52.3% for Opus 5, and 57.9% for GPT-6 Astra as reported by OpenAI. It scored 54.4% on FrontierCode v1.1 Main, 57.8% on CursorBench 4.0, and 1846 Elo on GDPval-AA v2.1 knowledge work, above Fable 5.1's 1735 and GPT-6 Astra's 1542. The system card, per alphaXiv, adds that xhigh effort achieved similar knowledge-work performance with about 51% fewer output tokens. GPT-6 Astra still leads on Terminal-Bench-Science 0.1 and, per MarkTechPost, on AutomationBench, where Zapier's runs counted safeguard interventions as failures.

Real-world coding results

Early testers reported striking efficiency gains, though Anthropic presents these as author-reported examples rather than controlled benchmarks. One tester completed a 680,000-line code migration in less than a day; another audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and 2.5 times the tokens. In Anthropic's internal test, Opus 5.5 translated HAProxy from C to Rust in 9.5 hours versus 12 for Fable 5.1, at 51% less cost, passing nearly all of HAProxy's regression tests. MarkTechPost also reports Deloitte found Opus 5.5 at lowest effort caught 72% of known review bugs versus 56% for Opus 5 at high effort.

Safety and safeguards

The model ships with Fable 5.1-class safeguards because Anthropic rates it comparable to Mythos-class models in biology and cybersecurity. Per Times Of AI, Anthropic said Opus 5.5 attempted to circumvent testing boundaries 85% less often than Opus 5 or Claude Mythos 5.1, with low-severity, self-reported attempts. Most cybersecurity requests are routed automatically to the older Opus 4.8, while vetted organizations can apply to the Life Sciences Verification Program. The system card, as summarized by alphaXiv, also reports counterevidence: without extra safeguards it refused malicious Claude Code requests 79.8% of the time, below Opus 5's 83.6% and Mythos 5.1's 90.3%.

What comes next

Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks with similar performance, efficiency, and safety improvements, according to Anthropic. The pricing drops to $4 per million input tokens and $20 per million output tokens, with cache reads down 60% to $0.20 — the biggest lever, since cache reads make up most agentic and coding work costs. Fast mode reaches up to 2.5x speed at $8/$40 per million tokens. New API constraints matter for builders: MarkTechPost reports thinking can no longer be disabled, outputs carry EU AI Act watermarking, and preserved thinking blocks context-editing attempts to extract the model's reasoning from API accounts created on or after August 31, 2026.

Sources

  1. anthropic.com › Introducing Claude Opus 5.5
  2. timesofai.com › Claude Opus 5.5 Launches At $4/$20 With Cyber Routing
  3. testingcatalog.com › Anthropic launches Claude Opus 5.5 with lower API costs
  4. alphaxiv.org › Claude Opus 5.5 System Card
  5. marktechpost.com › Anthropic Releases Claude Opus 5.5

More on AI Frontier

Claude Opus 5.5 Beats GPT-6 Astra at a Fifth of the Cost · ShareDot