The Shortcomings Of Astra Vs Fable’s Reduced Benchmark Criteria
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Shortcomings Of Astra Vs Fable’s Reduced Benchmark Criteria on ThorstenMeyerAI.com

TL;DR

Recent findings show Astra’s benchmark scores have shifted due to index revisions, and its architecture questions token-based efficiency metrics. Fable’s performance claims are based on outdated or inconsistent data, complicating direct comparisons.

Recent analysis indicates that Astra’s reported performance and efficiency advantages over Fable are based on outdated or inconsistent benchmark data, and that architectural differences significantly impact the validity of direct comparisons.

Thorsten Meyer, an AI researcher, examined Astra’s latest benchmark scores and architectural design, finding that the numbers cited in recent comparisons are based on different versions of the Artificial Analysis Index, which was revised shortly before Astra’s launch. This revision altered the scoring basket, leading to shifts in Astra’s scores from 66 to 55 in some evaluations, and from 61 to 54 in others, depending on the snapshot. Consequently, claims that Astra outperforms Fable on the Intelligence Index are based on numbers that are no longer current or consistent.

Further, Meyer highlights that the circulating narrative suggesting Astra “attacks the economics” of intelligence is misleading. The official Artificial Analysis report states Astra is approximately 75% more expensive than its predecessor, GPT-5.6 Sol, with a 2.5× increase in cost per task and only partial token-efficiency gains. The report concludes Astra is less cost-effective for general intelligence per dollar than the model it replaced. However, Astra does excel in coding tasks, where it achieves a genuine reduction in token usage and cost, positioning it on the Pareto frontier for coding agents.

Architectural differences further complicate comparisons. Astra employs a looped transformer architecture that reasons in latent space, which means it can process more tasks without emitting tokens or chain-of-thought explanations. This design renders token counts an unreliable proxy for compute and efficiency, as the index measures tokens, not the actual computational effort involved in Astra’s reasoning process. The model’s architecture externalizes much of its reasoning, making token-based metrics misleading when comparing it to models like Fable that rely on verbalized reasoning.

At a glance
analysisWhen: developing; recent benchmark revisions…
The developmentNew analysis reveals Astra’s benchmark scores and efficiency metrics are affected by index revisions and architectural differences, challenging previous claims about its performance and economics.
Five Points That Became Two — Reality Check
AI Dispatch · Reality Check · 5 September 2026

Five points that became two: what’s wrong with the Astra vs Fable benchmark

The comparison everyone is quoting — Fable 66, Astra 61, “not a rounding error” — is built on numbers that were stale when written, measuring a quantity that no longer means what it used to, aggregated in a way that hides the reversals that matter. The benchmark isn’t broken. The way it’s being read is.

Problem 1 — the numbers moved: Index v4.1.1 → v4.2 during launch week
as quotedAA today (v4.2)
Claude Fable 5.16657 max effort
GPT-6 Astra6155 max · a third source says 60
gap: 5 2 “Five points is not a rounding error.” Two points, on an aggregate of ten evals that just swapped three of them (GPQA Diamond out; AA-Briefcase + GDP.pdf in), is exactly a rounding error. Not AA’s fault — revising an index is how you keep it honest. The error is downstream: quote the version, or don’t quote the number.
◆ Problem 3 — the tell: Astra’s effort dial isn’t connected to the engine
non-reasoning554.4M tok · $990
medium52
xhigh54$2,778
max55$3,020
Non-reasoning = max. Same score, 3× the cost. Because Astra is reported to be a looped / recurrent-depth transformer — it reasons in latent space, without emitting tokens. The Index prices cost in tokens, measures verbosity in tokens, computes time in tokens. For this architecture it’s counting the receipt, not the work. “140M vs 42M tokens” compares Fable’s verbalized reasoning to Astra’s post-loop output — an artefact, not an efficiency finding. Nobody outside OpenAI knows what the loops cost in GPU-seconds.
The other three problems
02
AA’s own conclusion is the opposite of the story
AA’s benchmarking note: Astra is 75% more expensive than GPT-5.6 Sol at max effort and “largely sits behind its predecessor on the Intelligence Index vs cost frontier.” Price went 2.5× ($4/$20 → $10/$50); token savings only partly offset it. The genuine efficiency win lives in one place: the Coding Agent Index, where Astra equals Fable 5 at under half the cost. “Astra attacks the economics” stretched a true coding result over an intelligence index where AA says the reverse.
04
“Max effort” isn’t the same experiment twice
Fable at max = more tokens. Astra at max = ~nothing (see ladder). And OpenAI’s docs say Astra does not support `none` reasoning effort — yet AA lists a “non-reasoning” score. The most efficient-looking config on the leaderboard may not be one you can buy.
05
The aggregate hides the reversals
Index: Fable +2. OpenAI’s own evals (self-reported): Astra ahead 6 of 7 — AutomationBench, BenchCAD, Terminal-Bench 4.0, DeepSWE, TB-Science, FrontierMath T4; Fable takes HLE+tools. A 6–1 task split became a two-point average, and the average became the story. Ten choices deep, two points is noise wearing a number.
✓ What actually changed — and it’s not on the leaderboard
Hallucination rate 92% → 51% on AA-Omniscience — a 41-point drop; matters more than any 2 Index points “Same headline price” hides cache read $1.00 vs $0.25 (4×) + a 25% cache-write premium — the line that dominates agentic bills Coding Agent Index: Astra = Fable 5 at < half the cost — real, and narrow
The take

Three things happened at once: the Index was revised (five became two), the architecture changed (tokens stopped being compute), and the aggregate did what aggregates do (6–1 became +2). A leaderboard position now tells you less than it ever has — and the more advanced the architecture, the less it tells you. Latent reasoning is only the first architecture to break the token proxy. So with your Astra access: ignore the Index number. Take your ten real tasks. Run both models at the effort setting you’ll actually pay for. Measure the bill including the cache line. Measure the failure rate — the 41-point hallucination drop is the one number here I’d bet money on. The benchmark can’t decide for you anymore.

Sources: Artificial Analysis Intelligence Index v4.2 / v4.1.1 model pages (Fable 5.1; Astra non-reasoning/low/medium/high/xhigh/max — scores, tokens, total cost, cost-per-task method incl. cache-write & reasoning tokens) and “Benchmarking GPT-6 Astra” (75% vs Sol, cost frontier, Coding Agent Index, 92%→51% hallucination, 2.5× price, cache terms); OpenAI GPT-6 Astra developer docs (`none` unsupported, cache-write billing, logprobs removed); Alan D. Thompson, The Memo 4 Sep 2026 (looped-transformer read, unconfirmed); OpenAI’s self-reported Astra-vs-Fable table; the circulating 66/61 comparison (pre-v4.2). Scores are version-dependent and were changing at time of writing. Not investment advice.
thorstenmeyerai.com

Impact of Benchmark Revisions and Architecture on Performance Claims

This analysis underscores that performance and cost claims for Astra are heavily dependent on the specific benchmarks and architectures used. The shifting benchmark scores mean that comparisons based on outdated data can mislead stakeholders about Astra’s true capabilities and efficiency. The architectural differences, especially Astra’s latent-space reasoning, challenge the validity of token-based efficiency metrics, which have become a common standard for measuring AI performance. For users and developers, understanding these nuances is crucial for making informed decisions about model deployment and evaluation.

DULIWO Model Scriber Tool Kit, 7-Blade Chisel Set for Gunpla

DULIWO Model Scriber Tool Kit, 7-Blade Chisel Set for Gunpla

  • Complete Model Scriber Kit: Includes blades, drill bits, tweezers, and brush
  • High-Quality Materials: Tungsten steel blades and lightweight aluminum handle
  • Versatile Functionality: Engraving, cutting, scribing, burr removal, drilling, and cleaning

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Astra’s Benchmarking and Architectural Design

Thorsten Meyer’s recent review follows the release of Astra, a new GPT-6 variant, which was initially compared to Fable 5.1 based on an Artificial Analysis Index score of 66 versus 61. However, shortly before Astra’s launch, the index was revised from version 4.1.1 to 4.2, leading to significant score shifts across all models tested. This revision included dropping certain evaluation components and adding new ones, which altered the scoring basket and rendered previous comparisons outdated. Meanwhile, Astra’s architecture has been reported as a looped or recurrent-depth transformer, capable of reasoning in latent space without emitting tokens for every step, unlike traditional transformer models.

Prior to these developments, performance comparisons focused mainly on token counts and cost per task, with Fable claiming a significant advantage based on token efficiency. However, Meyer’s analysis suggests that these metrics are no longer valid indicators of true computational effort, especially given Astra’s architectural approach. The debate over what constitutes a fair comparison remains ongoing, with many experts emphasizing the importance of architecture-aware metrics.

“The numbers used in recent Astra vs. Fable comparisons are based on different versions of the benchmark index, which was revised shortly before Astra’s launch, making direct comparisons unreliable.”

— Thorsten Meyer

Amazon

AI performance analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Astra’s True Performance

It remains unclear how Astra’s architecture impacts real-world computational costs beyond token counts, as OpenAI has not disclosed the hardware or processing effort involved in its latent reasoning loops. The extent to which token efficiency translates into actual compute savings is still under investigation, and the impact of index revisions on past performance claims complicates longitudinal comparisons. Additionally, the broader implications for AI benchmarking standards and whether current metrics adequately reflect architectural differences are ongoing debates within the community.

Amazon

AI model efficiency testing devices

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in Benchmarking and Model Evaluation

Further independent analyses are expected as more data about Astra’s architecture and actual compute costs become available. OpenAI and third-party researchers are likely to develop new benchmarks that account for architectural differences, moving beyond token counts to measure real computational effort. Stakeholders should watch for updated performance reports, revised benchmarking standards, and potential clarifications from OpenAI regarding Astra’s architecture and efficiency metrics. These developments will help establish more accurate comparisons and inform deployment decisions.

Amazon

AI architecture comparison tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why are Astra’s benchmark scores changing?

The scores are shifting due to revisions in the Artificial Analysis Index, which updated the scoring methodology and evaluation components shortly before Astra’s launch, making previous scores outdated.

Does Astra really outperform Fable on intelligence per dollar?

According to the official Artificial Analysis report, Astra is less cost-effective than its predecessor for general intelligence tasks, though it excels in coding efficiency. The comparison depends heavily on the specific index and architecture considered.

How does Astra’s architecture affect efficiency measurements?

Astra’s architecture reasons in latent space with looped transformers, meaning token counts no longer accurately reflect the actual compute effort, challenging traditional token-based efficiency metrics.

What are the implications for AI benchmarking?

This analysis suggests that current benchmarking practices may be insufficient for architectural differences, indicating a need for new metrics that better reflect true computational effort.

Source: ThorstenMeyerAI.com

You May Also Like

Madrid Allocates 6 Million Euros To Promote AI Adoption In Companies

The Community of Madrid has announced a 6 million euro fund to support the integration of artificial intelligence in local businesses.

Startup Founders Urge U.S. Government Not To Shut Off Chinese Open Weight AI

Startup leaders appeal to U.S. government to avoid shutting down Chinese open-weight AI projects, citing innovation and global competitiveness concerns.

How ByteDance’s AI4S Investment Could Reshape STEM Careers

ByteDance launches the Seed STEM Scientist Program, seeking 100 researchers for a six-month AI-driven science pilot in Beijing, signaling new industry-academic ties.

The clause. How a contractual definition of AGI met the capital built on top of it.

Analysis of how the original AGI clause in the Microsoft-OpenAI contract was defused through amendments, shifting from a termination trigger to a verification process.