AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Five frontier AI models released in September 2026 — Claude Opus 5.5, Claude Fable 5.1, GPT-6 Astra, GPT-6 Sol and GPT-6 Luna — have been compared on the same independent yardstick, the Artificial Analysis Intelligence Index (v4.3). Claude Opus 5.5 posts the top score at 58, while GPT-6 Luna is by far the cheapest at roughly $0.07 per task. The comparison argues that cost per task, not raw intelligence ranking, is the deciding metric for buyers.

Five frontier AI models released within three weeks — Anthropic’s Claude Fable 5.1 and Claude Opus 5.5, and OpenAI’s GPT-6 Astra, GPT-6 Sol and GPT-6 Luna — have been compared on a single independent yardstick, and the results upend the launch-week rankings. On the current Artificial Analysis Intelligence Index (v4.3), Claude Opus 5.5 posts the highest score ever measured by the tracker at 58, while OpenAI’s budget-tier GPT-6 Luna completes indexed tasks for roughly $0.07 each — a spread of nearly 100x in cost per task between the cheapest and most capable models. The comparison’s central finding: ranking models by intelligence alone is the wrong question for buyers; the useful measure is which model clears a quality bar at the lowest cost per task.

The September release calendar was unusually dense. Anthropic shipped Claude Fable 5.1 on September 1 and Claude Opus 5.5 on September 22. OpenAI released GPT-6 Astra on September 3, then GPT-6 Sol and GPT-6 Luna on the same day as Opus 5.5. All five are now measured on the same independent index, allowing a direct score-versus-cost comparison at matched effort levels.

At maximum effort, the results break into three tiers. Opus 5.5 leads on score at 58, with per-token pricing of $4 input / $20 output per million tokens and a cost per indexed task of $5.98. GPT-6 Astra and Fable 5.1 tie at roughly 53, but at very different economics: Astra costs $3.26 per task at $10/$50 pricing, while Fable 5.1 costs $7.63 — now outscored by its cheaper Anthropic sibling. GPT-6 Sol (score 48, $1.06 per task) and GPT-6 Luna (score 37, $0.07 per task) occupy the budget end of the frontier.

The per-task cost picture sharpens the differences. For $100 at maximum effort, the comparison estimates roughly 17 tasks for Opus 5.5, 13 for Fable 5.1, 31 for Astra, 94 for Sol and 1,429 for Luna. Notably, Opus 5.5’s high per-task cost despite mid-tier token pricing is driven by verbosity: at max effort it writes roughly 119,000 output tokens per task, versus about 27,000 for GPT-6 Astra, according to the comparison.

At a glance
analysisWhen: published late September 2026, covering…
The developmentA new side-by-side comparison of all five frontier AI models released in September 2026 measures them on the re-based Artificial Analysis Intelligence Index (v4.3) and on cost per task.
Comparing Five Frontier AI Models In One Bill

Frontier AI · September 2026 · AA Intelligence Index v4.3

Comparing Five Frontier AI Models in One Bill

Claude Opus 5.5, Claude Fable 5.1, GPT-6 Astra, GPT-6 Sol and GPT-6 Luna — five frontier models released within three weeks, finally measured on one independent yardstick. The verdict: cost per task, not intelligence ranking, is the metric that decides procurement.

Top Score Ever · Opus 5.5 58
Cheapest Task · GPT-6 Luna $0.07

“Ranking them by intelligence is the wrong question. The useful one is which one clears your quality bar at the lowest cost per task.”

— The comparison (Thorsten Meyer AI)
5
Frontier Models
~100×
Cost Spread Per Task
1,429
Luna Tasks for $100
119k
Opus Output Tokens / Task
1M
Opus & Luna Context

01 · The Scoreboard

Three Tiers at Maximum Effort

At matched effort on the re-based index, the field splits into a clear leader, a tied middle, and a budget pair. Intelligence scores tell only half the story — the per-task economics tell the other half.

Tier 1 · Flagship

Claude Opus 5.5

58 INDEX
$/task$5.98
Price /M in–out$4 / $20
Context1M
Tier 2 · Contenders

GPT-6 Astra

53 INDEX
$/task$3.26
Price /M in–out$10 / $50
Context1.05M
Tier 2 · Contenders

Claude Fable 5.1

53 INDEX
$/task$7.63
Effort measuredxhigh / max
Contextunlisted
Tier 3 · Budget

GPT-6 Sol

48 INDEX
$/task$1.06
Effort measuredlow / max
Context872k
Tier 3 · Budget

GPT-6 Luna

37 INDEX
$/task$0.07
Price /M in–out$0.10 / $0.50
Context1M

02 · The Economics

What $100 Buys at Maximum Effort

Task throughput per $100 exposes a spread of roughly 100× between the cheapest and most capable models — and reframes “best” as a workload question, not a leaderboard question.

GPT-6 LUNAscore 37
1,429 TASKS
GPT-6 SOLscore 48
94 TASKS
GPT-6 ASTRAscore 53
31 TASKS
CLAUDE OPUS 5.5score 58
17 TASKS
CLAUDE FABLE 5.1score 53
13 TASKS

⚠ Opus 5.5’s high per-task cost despite mid-tier token pricing is driven by verbosity: ~119,000 output tokens per task vs ~27,000 for GPT-6 Astra.

03 · The Measurement Shift

The Re-Based Index That Changed the Scores

Artificial Analysis rebuilt its index in September 2026. Fable 5.1’s launch-week 66 on the old v4.1.1 scale became a 53 on the current v4.3 scale. “The model didn’t get worse. The test got harder.” Mixing launch-week headlines means comparing different rulers.

FABLE v4.1.1
66
FABLE v4.3
53
OPUS v4.3
58
ASTRA v4.3
53
LUNA v4.3
37

All figures cited: current v4.3 index. Maximum effort ≠ default — real deployments at medium/high effort may shift rankings and costs.

04 · Trade-Offs

Where Each Model Wins — and Where It Doesn’t

ModelStrengthValue PositionFlagged Weakness
Claude Opus 5.5 Top score 58; 1822 Elo on AA-Briefcase knowledge work, 143 pts ahead of Fable 5.1; first Anthropic model to beat GPT-5.6 Sol on presentation quality Best score, premium cost ✗ ~119k output tokens/task at max effort; cybersecurity tasks re-route via fallback to older Claude models
GPT-6 Astra Ties Fable at 53; strong token efficiency (~27k output tokens/task) Best value near the top ~ Not the outright score leader
Claude Fable 5.1 Safeguards suited to workloads where buyers’ own tests show it ahead ✗ Outscored by a cheaper sibling at higher cost — the model to question Context window unlisted; measured only at xhigh/max effort
GPT-6 Sol Score 48 at $1.06/task — balanced mid-budget option ✓ Strong cost-to-capability ratio ✗ Regressed on polished, complete deliverables (GDPval-AA)
GPT-6 Luna Sub-cent-to-seven-cent tasks — ideal for high-volume classification, extraction, routing Cheapest frontier option ✗ Coding Agent Index slipped 2 points vs predecessor

05 · Buyer Playbook

The Shadow-Test Path to Procurement

Benchmark scores are historical measurements on specific evaluations — not guarantees on your workload. The comparison’s recommended next step is a shadow test before any switch.

1

📋 Sample Your Tasks

Pull a representative sample of real tasks from your own production workload — not benchmark prompts.

2

⚖️ Run Shadow Tests

Run each candidate model on the same sample; compare quality and cost directly, since index results may not transfer.

3

🎚️ Tune Effort Level

Opus 5.5: start at medium effort, escalate only when needed. Astra: test computer-use and agent workloads where token efficiency pays off.

4

🔍 Verify the Premium

Fable 5.1 buyers: confirm with your own tests that it beats Opus 5.5 on your tasks before paying the premium.

5

🔄 Re-Check the Index

Further AA revisions could shift numbers again — validate rankings against the current version before procurement.

06 · Fine Print

Caveats Buyers Should Weigh

Incomplete Spec Sheet

Fable 5.1’s context window is unlisted in the sources used for this comparison.

Re-Based Mid-Month

Scores published before the v4.3 change — including Fable 5.1’s launch-week 66 — are not comparable with current figures.

Partial Effort Curves

Fable 5.1 measured only at xhigh/max; Sol and Luna only at low/max. Intermediate cost-versus-score curves are partially estimated.

Fallback & Verbosity Unknowns

How Anthropic’s fallback mechanism affects real-world security and biology work beyond index scores — and whether max-effort verbosity persists at medium/high effort — remain open questions.

Why Cost per Task Beats Intelligence Rankings

The comparison matters because buyers face a practical decision, not a leaderboard. A team running high-volume classification, extraction or routing can use GPT-6 Luna at sub-cent-to-seven-cent task costs — roughly two orders of magnitude cheaper than flagship models — if its score of 37 meets their quality bar. Teams producing client-ready knowledge work or long agentic coding get the strongest independent score from Opus 5.5, which posted 1822 Elo on the private AA-Briefcase knowledge-work evaluation, 143 points ahead of Fable 5.1 and, per the comparison, the first Anthropic model to beat GPT-5.6 Sol on presentation quality.

The value positioning is also shifting. GPT-6 Astra ties Fable 5.1 on score at well under half the cost per task, making it, in the comparison’s framing, the best value near the top of the table. Fable 5.1 emerges as the model to question: it is now outscored by a cheaper sibling at a higher cost per task, though Anthropic positions it with safeguards suited to workloads where buyers’ own tests show it ahead.

The comparison also flags model-specific trade-offs: most cybersecurity tasks on Anthropic models re-route via fallback to older Claude models, which can depress scores on security and biology work; GPT-6 Sol regressed on polished, complete deliverables (GDPval-AA); and GPT-6 Luna’s Coding Agent Index slipped 2 points versus its predecessor.

Amazon

AI model cost analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Re-Based Index That Changed the Scores

The most consequential background detail is that Artificial Analysis rebuilt its index this month. When Fable 5.1 launched on September 1, it scored 66 on the older v4.1.1 version of the index — the highest score ever recorded at the time. On the current v4.3 version, the same model scores 53, tied with GPT-6 Astra. The model did not get worse; the test got harder. Any comparison mixing launch-week headlines from different weeks is, as the analysis puts it, comparing different rulers. All figures cited here come from the current v4.3 index.

Two further measurement caveats apply. First, Anthropic models run with fallback behavior: safety-flagged requests route to older Claude models, which can lower scores on security and biology tasks. Second, maximum effort is not the default — most real deployments run at medium or high effort, where rankings and per-task costs can differ from the max-effort figures shown. Fable 5.1’s context window was not listed in the sources used for the comparison, while Opus 5.5 supports 1M tokens, Astra 1.05M, Luna 1M and Sol 872k.

“Ranking them by intelligence is the wrong question. The useful one is which one clears your quality bar at the lowest cost per task.”

— The comparison (Thorsten Meyer AI)

Amazon

AI task management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Caveats Buyers Should Weigh

Several limits remain. Fable 5.1’s context window is unlisted in the sources used, leaving an incomplete spec sheet. The index was re-based mid-month, so scores published before the v4.3 change — including Fable 5.1’s launch-week 66 — are not comparable with current figures. Fable 5.1 was measured only at xhigh and max effort, and Sol and Luna only at low and max, so the cost-versus-score curves for intermediate effort levels are partially estimated.

It is also unclear how the Anthropic fallback mechanism affects real-world performance on security and biology tasks beyond its measured effect on index scores, and whether max-effort verbosity (such as Opus 5.5’s ~119k output tokens per task) persists at the medium and high effort levels most deployments actually use. Benchmark scores are historical measurements on specific evaluations, not guarantees of performance on any buyer’s workload.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing on Your Own Workload

The comparison’s recommended next step for buyers is a shadow test: run each candidate model on a sample of real tasks from your own workload and compare quality and cost directly, since index results may not transfer. For teams considering Opus 5.5, the guidance is to run at medium effort first and escalate only when needed, given the token volume at max effort. Buyers weighing Astra should evaluate it for computer-use and agent workloads where its token efficiency pays off; those weighing Fable 5.1 should verify with their own tests whether it outperforms Opus 5.5 on their tasks before paying the premium. Further index revisions from Artificial Analysis — which re-based the scale this month — could again shift the numbers, so rankings should be re-checked against the current version before any procurement decision.

Amazon

AI model pricing calculators

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Which of the five models scored highest?

Claude Opus 5.5 scored 58 on the Artificial Analysis Intelligence Index (v4.3) at maximum effort — the highest score the tracker has measured, by several points, on the current version of the index.

Which model is the cheapest to run?

GPT-6 Luna, priced at $0.10 input / $0.50 output per million tokens, completes an indexed task for roughly $0.07. For $100 at max effort, the comparison estimates about 1,429 tasks, versus 17 for Opus 5.5.

Why did Claude Fable 5.1’s score drop from 66 to 53?

Artificial Analysis re-based its index from v4.1.1 to v4.3 during September 2026. The same model scored 66 on the old version and 53 on the new one — the model did not change, the test became harder. Scores should only be compared within a single index version.

Why does Opus 5.5 cost more per task than its token price suggests?

At maximum effort, Opus 5.5 writes roughly 119,000 output tokens per task, versus about 27,000 for GPT-6 Astra. Despite less than half of Astra’s per-token price, that verbosity pushes its cost per task to $5.98 — nearly twice Astra’s $3.26.

Should I pick a model based on these benchmark scores alone?

No. The comparison advises running a shadow test on your own tasks before switching. Most deployments run at medium or high effort rather than max, where rankings and costs differ, and Anthropic’s fallback routing to older models can affect results on certain task types.

Source: Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Oliver Blume’s VW Jobs Victory Leaves Power Questions Untouched

Oliver Blume secures VW jobs, but key questions about corporate influence and leadership remain unresolved amid rising interest in VW’s future leadership.

Outcome-First Decisions: Keep, Change, or Kill

Thorsten Meyer AI published an open-source framework for portfolio reviews that assigns keep, change or kill verdicts.

The Future Of Mobile Workstations: Top 9 AI-Driven Laptops For 2026

A 2026 comparison ranks nine mobile workstations, led by Dell’s Precision 7680, while exposing major graphics and battery tradeoffs.

10 Best Soundbars In 2026

A 2026 comparison ranks 10 soundbars by audio formats, connectivity, power, room fit and setup, with Sonos, Bose and Samsung among the picks.