The AI Race Heats Up: Qwen3.8-Max's Latest Numbers In Focus

📊 Full opportunity report: The AI Race Heats Up: Qwen3.8-Max's Latest Numbers In Focus on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has publicly confirmed its new flagship model, Qwen3.8-Max, featuring 2.4 trillion parameters and top-tier benchmark performance. The open weights are set to ship next week, signaling a major development in the AI race.

Alibaba has officially confirmed the specifications and benchmark performance of Qwen3.8-Max, its largest-ever AI model with 2.4 trillion parameters. This marks a significant milestone in the AI model race, as the company prepares to release open weights next week, intensifying competition among leading AI developers.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, revealing a model built on the Qwen3.5 architecture with 95 billion active parameters per query via sparse mixture-of-experts technology. The model demonstrated strong results on several benchmarks, including a Terminal-Bench score of 86.6, surpassing models like Claude Fable 5 and only behind GPT-5.6 Sol at maximum effort.

Alibaba disclosed that the model is multimodal, capable of processing text, images, and videos, with a context window of 983,616 tokens. The benchmark results also show significant improvements in agentic and long-horizon tasks, with the model outperforming its predecessor in areas like deep software engineering benchmarks and long-term reasoning, notably tripling scores in some agentic tasks.

While the flagship model’s performance is impressive, Alibaba also announced a smaller, more deployable version, Qwen3.8-27B, which is designed for local deployment on high-memory machines. The open weights for this smaller model are scheduled to be released next week, emphasizing Alibaba’s focus on broad accessibility and practical deployment.

At a glance
updateWhen: announced August 3, 2023; benchmarks an…
The developmentAlibaba announced the official release and benchmark results of Qwen3.8-Max, confirming its 2.4 trillion parameters and open weights scheduled for next week.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Impact of Alibaba’s Open Model Release on AI Competition

The official confirmation of Qwen3.8-Max's specifications and benchmark results signifies a major advancement in the AI model landscape, especially with the upcoming release of open weights. This development could shift the competitive dynamics by providing researchers and developers access to a state-of-the-art, multimodal model with substantial agentic capabilities. It also raises questions about how this will influence the market share of existing models and the pace of AI innovation, given Alibaba’s strategic emphasis on transparency and openness.

Amazon

AI development hardware high-memory servers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Recent Developments in Large-Scale AI Models

Over the past two weeks, the AI community has been tracking Alibaba’s stealthy development of Qwen3.8-Max, which was first hinted at through an anonymous model named 'kaleb' on the Code Arena leaderboard and later confirmed during the World AI Conference in Shanghai. The model's preview, initially available through a limited preview endpoint at discounted pricing, generated significant interest due to its size and capabilities. Meanwhile, competitors like Moonshot’s Kimi K3 and other major players have been releasing their models, intensifying the race for the largest and most capable AI models.

Alibaba’s approach has involved strategic timing, revealing the full benchmark table after a two-week period of anticipation. The model’s architecture builds on prior Qwen versions, emphasizing sparse mixture-of-experts for scaling, and its multimodal capabilities position it as a versatile competitor in the AI space. The upcoming open weights are seen as a move toward broader adoption and deployment, especially for the 27B version optimized for local use.

"We are excited to share the full benchmark results and will release open weights next week to foster innovation and accessibility in AI development."

— Alibaba spokesperson

Building Production-ready Applications With Large Language Models Handbook: From Foundation Models to Scalable AI Systems Using Modern LLM ... Enterprise Tools for Real-World Deployment

Building Production-ready Applications With Large Language Models Handbook: From Foundation Models to Scalable AI Systems Using Modern LLM ... Enterprise Tools for Real-World Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unconfirmed Aspects of the Open-Weight Release and Licensing

Details about the licensing terms for the 2.4 trillion-parameter weights remain unpublished, raising questions about usage rights and restrictions. It is also unclear whether the open weights will include the full model or a pruned version optimized for deployment. Additionally, the long-term performance of the smaller Qwen3.8-27B model, especially regarding agentic capabilities after compression, has yet to be validated through independent benchmarks.

Amazon

AI model training GPU clusters

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s Open-Source AI Strategy

Next week, Alibaba will release the open weights for Qwen3.8-27B, enabling local deployment and testing by third parties. The company is expected to publish detailed licensing terms and possibly further benchmarks for the smaller model. Industry observers will closely monitor how the open weights impact adoption, competitive dynamics, and whether Alibaba’s claims about agentic improvements hold up in real-world applications.

Amazon

multimodal AI processing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the key specifications of Alibaba’s Qwen3.8-Max?

Qwen3.8-Max features 2.4 trillion parameters, built on the Qwen3.5 architecture, with 95 billion active parameters per query, and supports multimodal inputs including text, images, and videos.

When will the open weights for Qwen3.8-27B be available?

Alibaba plans to release the open weights for Qwen3.8-27B next week, making the smaller, deployable version accessible for local use.

How does Qwen3.8-Max compare to other leading models?

In benchmark tests, Qwen3.8-Max outperforms models like Claude Fable 5 and is only behind GPT-5.6 Sol at maximum effort, showing competitive performance in several tasks.

What are the implications of Alibaba’s open-weight strategy?

The open release could democratize access to large-scale multimodal models, accelerate research, and challenge existing market leaders, depending on licensing and deployment terms.

What remains uncertain about Alibaba’s new model?

Key uncertainties include licensing details, the full capabilities of the 27B model after compression, and how agentic performance will translate into practical applications.

Source: ThorstenMeyerAI.com

You May Also Like

The Nordics: Protect the Worker, Not the Job

Exploring how Nordic countries prioritize worker security over job preservation, enabling smoother transitions amid automation and economic shifts.

Monitoring Tech Operations: Apple’s Strategic Response To Trade Secret Theft

Apple has filed a lawsuit against OpenAI, alleging former employees stole trade secrets. This highlights increased corporate efforts to monitor tech espionage.

Glasspane: One Dataset, Three Views

Glasspane introduces a new approach to infrastructure monitoring with a single dataset presented through role-specific views, emphasizing trust and transparency.

Aleph Alpha. The retrospective case.

Analyzing Aleph Alpha’s strategic pivot, acquisition, and what it reveals about Europe’s AI capabilities and timing challenges.