Five Points That Became Two: What’s Wrong With The Astra Vs Fable Benchmark
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Five Points That Became Two: What’s Wrong With The Astra Vs Fable Benchmark on ThorstenMeyerAI.com

TL;DR

A recent benchmark comparing GPT-6 Astra and Fable 5.1 is flawed due to index revisions, architectural differences, and misinterpretations of token efficiency. The actual performance and economics are more nuanced than the headline figures suggest.

Recent claims that GPT-6 Astra outperforms Fable 5.1 on the Artificial Analysis Intelligence Index are based on outdated and inconsistent benchmark data, according to a detailed analysis by Thorsten Meyer. The findings challenge the narrative that Astra’s architecture offers superior intelligence per dollar, highlighting fundamental flaws in the benchmarking approach and interpretation.

Thorsten Meyer’s investigation shows that the widely circulated comparison between Astra and Fable is anchored in benchmark scores that have since been revised. The original figures claimed a five-point lead for Fable (66 vs. 61), but recent data indicates the scores are now closer—Fable at 57 and Astra at 55—due to index updates and re-scoring across different evaluation versions. This discrepancy underscores how benchmark revisions can distort perceived performance gaps.

Furthermore, the narrative that Astra ‘attacks the economics’ of AI intelligence is misleading. While Astra demonstrates efficiency in coding tasks—achieving comparable performance at less than half the cost of Fable—the same does not hold for general intelligence metrics. According to Artificial Analysis, Astra is actually less cost-effective than its predecessor, GPT-5.6 Sol, on the broader Intelligence Index, with its price increasing 2.5 times and only partial token-efficiency gains.

Adding to the confusion, Astra’s architectural design—featuring latent reasoning loops—renders token-based efficiency metrics obsolete for measuring true compute costs. The Index relies heavily on token counts, which do not accurately reflect the model’s reasoning process, especially when reasoning occurs outside of token generation. As a result, comparing token usage between Astra and Fable is misleading, conflating architectural differences with performance and efficiency.

At a glance
analysisWhen: published April 2024
The developmentNew analysis from Thorsten Meyer reveals that the Astra vs Fable benchmark is based on outdated and misinterpreted data, leading to misleading conclusions about model performance and cost-efficiency.
Five Points That Became Two — Reality Check
AI Dispatch · Reality Check · 5 September 2026

Five points that became two: what’s wrong with the Astra vs Fable benchmark

The comparison everyone is quoting — Fable 66, Astra 61, “not a rounding error” — is built on numbers that were stale when written, measuring a quantity that no longer means what it used to, aggregated in a way that hides the reversals that matter. The benchmark isn’t broken. The way it’s being read is.

Problem 1 — the numbers moved: Index v4.1.1 → v4.2 during launch week
as quotedAA today (v4.2)
Claude Fable 5.16657 max effort
GPT-6 Astra6155 max · a third source says 60
gap: 5 2 “Five points is not a rounding error.” Two points, on an aggregate of ten evals that just swapped three of them (GPQA Diamond out; AA-Briefcase + GDP.pdf in), is exactly a rounding error. Not AA’s fault — revising an index is how you keep it honest. The error is downstream: quote the version, or don’t quote the number.
◆ Problem 3 — the tell: Astra’s effort dial isn’t connected to the engine
non-reasoning554.4M tok · $990
medium52
xhigh54$2,778
max55$3,020
Non-reasoning = max. Same score, 3× the cost. Because Astra is reported to be a looped / recurrent-depth transformer — it reasons in latent space, without emitting tokens. The Index prices cost in tokens, measures verbosity in tokens, computes time in tokens. For this architecture it’s counting the receipt, not the work. “140M vs 42M tokens” compares Fable’s verbalized reasoning to Astra’s post-loop output — an artefact, not an efficiency finding. Nobody outside OpenAI knows what the loops cost in GPU-seconds.
The other three problems
02
AA’s own conclusion is the opposite of the story
AA’s benchmarking note: Astra is 75% more expensive than GPT-5.6 Sol at max effort and “largely sits behind its predecessor on the Intelligence Index vs cost frontier.” Price went 2.5× ($4/$20 → $10/$50); token savings only partly offset it. The genuine efficiency win lives in one place: the Coding Agent Index, where Astra equals Fable 5 at under half the cost. “Astra attacks the economics” stretched a true coding result over an intelligence index where AA says the reverse.
04
“Max effort” isn’t the same experiment twice
Fable at max = more tokens. Astra at max = ~nothing (see ladder). And OpenAI’s docs say Astra does not support `none` reasoning effort — yet AA lists a “non-reasoning” score. The most efficient-looking config on the leaderboard may not be one you can buy.
05
The aggregate hides the reversals
Index: Fable +2. OpenAI’s own evals (self-reported): Astra ahead 6 of 7 — AutomationBench, BenchCAD, Terminal-Bench 4.0, DeepSWE, TB-Science, FrontierMath T4; Fable takes HLE+tools. A 6–1 task split became a two-point average, and the average became the story. Ten choices deep, two points is noise wearing a number.
✓ What actually changed — and it’s not on the leaderboard
Hallucination rate 92% → 51% on AA-Omniscience — a 41-point drop; matters more than any 2 Index points “Same headline price” hides cache read $1.00 vs $0.25 (4×) + a 25% cache-write premium — the line that dominates agentic bills Coding Agent Index: Astra = Fable 5 at < half the cost — real, and narrow
The take

Three things happened at once: the Index was revised (five became two), the architecture changed (tokens stopped being compute), and the aggregate did what aggregates do (6–1 became +2). A leaderboard position now tells you less than it ever has — and the more advanced the architecture, the less it tells you. Latent reasoning is only the first architecture to break the token proxy. So with your Astra access: ignore the Index number. Take your ten real tasks. Run both models at the effort setting you’ll actually pay for. Measure the bill including the cache line. Measure the failure rate — the 41-point hallucination drop is the one number here I’d bet money on. The benchmark can’t decide for you anymore.

Sources: Artificial Analysis Intelligence Index v4.2 / v4.1.1 model pages (Fable 5.1; Astra non-reasoning/low/medium/high/xhigh/max — scores, tokens, total cost, cost-per-task method incl. cache-write & reasoning tokens) and “Benchmarking GPT-6 Astra” (75% vs Sol, cost frontier, Coding Agent Index, 92%→51% hallucination, 2.5× price, cache terms); OpenAI GPT-6 Astra developer docs (`none` unsupported, cache-write billing, logprobs removed); Alan D. Thompson, The Memo 4 Sep 2026 (looped-transformer read, unconfirmed); OpenAI’s self-reported Astra-vs-Fable table; the circulating 66/61 comparison (pre-v4.2). Scores are version-dependent and were changing at time of writing. Not investment advice.
thorstenmeyerai.com

Implications of Benchmark Flaws on AI Performance Claims

This analysis highlights that the widely accepted Astra vs. Fable benchmark is fundamentally flawed, impacting how AI performance and economics are understood. Relying on outdated or misinterpreted data can lead developers and users to overestimate Astra’s efficiency and underestimate its costs, potentially skewing investment and development priorities. The findings emphasize the importance of transparent, architecture-aware benchmarking practices to accurately assess AI capabilities and value.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of the Astra vs Fable Benchmark Dispute

The comparison between Astra and Fable gained prominence after initial reports claimed Astra’s superior cost-efficiency and performance on the Artificial Analysis Intelligence Index. These claims were based on a benchmark score that circulated widely but was derived from an earlier version of the index. Over time, the index was revised—version 4.1.1 to 4.2—leading to recalibrated scores for all models involved. Meanwhile, Astra’s architectural innovations, such as latent reasoning loops, have complicated token-based efficiency measurements, which are central to the original benchmarking approach.

Initially, the narrative suggested Astra was a breakthrough in intelligence per dollar, driven by token efficiency. However, subsequent analysis revealed that the scoring metrics were based on different evaluation versions and architectures, making direct comparisons unreliable. Experts like Sebastian Raschka and Alan Thompson have pointed out that Astra’s architecture fundamentally changes how its performance should be measured, especially regarding compute costs and reasoning efficiency.

“The benchmark scores are not only outdated but also misrepresent the true performance of Astra and Fable, mainly because of index revisions and architectural differences.”

— Thorsten Meyer

Unresolved Questions About Astra’s True Performance

It remains unclear how Astra’s architectural features—particularly its latent reasoning loops—will impact real-world performance and cost-efficiency outside of token-based metrics. OpenAI has not publicly disclosed detailed cost data for the latent reasoning process, and external observers cannot verify the actual compute costs involved. Additionally, the long-term implications of index revisions and their influence on model rankings are still being evaluated, leaving some uncertainty about Astra’s true standing in the AI landscape.

Next Steps for Accurate Benchmarking and Model Evaluation

Developers and researchers are expected to focus on architecture-aware benchmarks that account for Astra’s latent reasoning capabilities and true compute costs. OpenAI and other organizations may update their evaluation protocols to better reflect these architectural innovations. Meanwhile, independent analysts will likely scrutinize Astra’s performance across diverse tasks, beyond token counts, to establish a clearer picture of its real-world efficiency and capabilities. Transparency from OpenAI regarding the architecture and operational costs will be crucial for an accurate assessment.

Key Questions

Why is the Astra vs Fable benchmark considered flawed?

The benchmark is flawed because it relies on outdated index versions, conflates architectural differences with performance metrics, and uses token counts that no longer accurately reflect Astra’s compute costs due to its latent reasoning loops.

Does Astra outperform Fable in terms of intelligence?

According to the latest data, Astra is not necessarily more intelligent per dollar in general tasks. It shows efficiency gains in coding but is less cost-effective on broader intelligence metrics, especially after index revisions.

What are latent reasoning loops, and why do they matter?

Latent reasoning loops are architectural features allowing Astra to reason internally without emitting tokens for every step. They make token-based efficiency metrics misleading because they externalize less of the model’s work, obscuring true compute costs.

Will the benchmark be revised to reflect Astra’s architecture?

It is likely that future benchmarks will incorporate architecture-aware metrics that better capture Astra’s latent reasoning capabilities, but current standards still rely heavily on token counts.

How should consumers interpret Astra’s performance claims?

Consumers should be cautious and consider that current performance and efficiency claims may be based on outdated or incomplete metrics. A deeper understanding of architectural differences is necessary for accurate assessment.

Source: ThorstenMeyerAI.com

You May Also Like

15 Best Propane Patio Heaters for Cozy Outdoor Spaces in 2025

An array of top propane patio heaters for cozy outdoor spaces in 2025 awaits, offering unmatched warmth and comfort you won’t want to miss.

60% of Consumers Are Reported to Engage With Voice Assistants, Study Finds

I discovered that 60% of consumers actively use voice assistants, hinting at a transformative shift in how we communicate and shop. What comes next?

12 Best Hydration Packs for Adventure Seekers in 2026

Discover the top hydration packs of 2026. Find the best overall, budget-friendly options, and specialized packs for every adventure in this expert roundup.

FreeOberon – Open-Source, Cross-Platform, Free Pascal/Turbo Pascal-Like Language

FreeOberon is an open-source, cross-platform IDE and compiler for the Oberon language, inspired by Pascal and Modula-2, available for Windows, macOS, and Linux.