The Numbers Behind Qwen3.8-Max’s AI Capabilities: What’s The Truth?

📊 Full opportunity report: The Numbers Behind Qwen3.8-Max’s AI Capabilities: What’s The Truth? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has confirmed the specifications of Qwen3.8-Max, including its 2.4 trillion total parameters and 95 billion active parameters. The model’s benchmark results show strong performance, but only on selected tests. Open weights will be released next week, with a smaller, deployable 27B version also coming.

Alibaba has officially confirmed the specifications of its flagship AI model, Qwen3.8-Max, revealing it has 2.4 trillion total parameters and performs strongly on several benchmarks. This marks a significant milestone in large language model development and opens the door for broader access to high-capacity AI models.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, previously previewed in July. The model is built on the Qwen3.5 architecture, employs sparse mixture-of-experts, and features a 95 billion active-parameter count, indicating that only about 4% of its total parameters are active per token. The model supports multimodal inputs—text, images, videos—and outputs text.

The benchmark results show the model outperforming several competitors on key tests: it scores 86.6 on Terminal-Bench 2.1, surpassing Claude models and only trailing GPT-5.6. It tops the PaperBench at 93.0 and demonstrates notable strength in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5. However, it trails significantly on deep software engineering benchmarks like SWE-bench Pro, where it scores 67.7 compared to Fable 5’s 80.0, highlighting areas for improvement.

The open weights for the model’s 2.4 trillion parameters are scheduled for release next week, although the current accessible version is a 95 billion-parameter subset optimized for deployment. Alibaba also announced a smaller, 27-billion-parameter version, Qwen3.8-27B, suitable for local deployment and running on a single high-memory machine, with details to be published soon.

At a glance
updateWhen: announced August 3, 2023; full details…
The developmentAlibaba announced the full specifications and benchmark results for Qwen3.8-Max, confirming its 2.4 trillion parameters and open weights release scheduled for next week.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Model Specifications and Benchmark Results

The confirmation of Qwen3.8-Max’s parameters and benchmark performance underscores Alibaba’s advancing position in large language models, especially in multimodal and agentic tasks. The open release of weights next week could impact AI deployment strategies, enabling wider access to high-capacity models and fostering innovation in AI applications. However, the model’s limitations on deep software engineering benchmarks suggest ongoing challenges in specific technical domains.

For developers and researchers, the availability of a 95 billion active-parameter model with open weights offers new opportunities for customization and integration, though practical deployment remains complex due to hardware requirements. The smaller 27B model aims to bridge that gap, providing a more accessible option for local inference and daily use.

Acer Aspire Go 15 AI Ready Laptop | 15.6" FHD (1920 x 1080) IPS Display | AMD Ryzen 7 7730U | AMD Radeon Graphics | 16GB DDR4 | 512GB PCIe Gen4 SSD | Wi-Fi 6 | Windows 11 Home | AG15-42P-R9FW

Acer Aspire Go 15 AI Ready Laptop | 15.6" FHD (1920 x 1080) IPS Display | AMD Ryzen 7 7730U | AMD Radeon Graphics | 16GB DDR4 | 512GB PCIe Gen4 SSD | Wi-Fi 6 | Windows 11 Home | AG15-42P-R9FW

  • Powerful AMD Ryzen 7 Processor: Smooth, responsive performance for multitasking
  • 16GB DDR4 Memory: Enhanced multitasking and faster data access
  • 512GB PCIe Gen4 SSD: Fast storage for quick file access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Development Timeline of Alibaba’s AI Models

Alibaba’s AI model journey has been marked by strategic previews and selective disclosures. In July, the company previewed Qwen3.8-Max during the World AI Conference in Shanghai, initially without detailed specifications. The model was identified as “kaleb” on the Code Arena leaderboard, and Alibaba confirmed its identity shortly after. The model’s capabilities and benchmark performance were kept under wraps until today’s full disclosure.

Prior to Qwen3.8-Max, Alibaba released smaller models like Kimi K3 and has steadily increased model size and complexity, emphasizing multimodal and agentic capabilities. The recent release aligns with broader industry trends toward larger, more capable models with open weights, but with a cautious approach to transparency and licensing.

The recent performance claims and benchmark results are part of a broader strategy to establish Alibaba as a leading AI innovator, competing with US and European firms, while managing the risks associated with large-scale model deployment.

"We are committed to open weights and advancing AI accessibility, with the full release scheduled for next week."

— Alibaba spokesperson

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Deployment and Capabilities

It is still unclear how the open weights will be licensed and whether they will be fully open-source or subject to restrictions. The practical implications of deploying a 2.4 trillion-parameter model remain uncertain, given hardware limitations and infrastructure requirements. Additionally, the performance of the smaller 27B version in real-world scenarios is yet to be demonstrated, and whether it can retain the agentic improvements seen in the flagship model is still unknown.

Further details about licensing terms, deployment options, and long-term performance are expected in the coming weeks.

Design Beyond Devices: Creating Multimodal, Cross-Device Experiences

Design Beyond Devices: Creating Multimodal, Cross-Device Experiences

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s AI Model Access and Adoption

Next week, Alibaba plans to release the full 2.4 trillion-parameter model weights, enabling researchers and developers to experiment with the model directly. The smaller Qwen3.8-27B will also become available for local deployment, offering a practical option for daily AI tasks. Industry observers will closely monitor how these models perform in diverse applications and whether licensing terms facilitate widespread adoption.

Further benchmark results and technical documentation are anticipated, providing clarity on the model’s capabilities and limitations, and shaping future AI development strategies.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main capabilities of Qwen3.8-Max?

Qwen3.8-Max features multimodal input support (text, images, videos), strong benchmark performance in tasks like multimodal reasoning and agentic execution, and a 95-billion active-parameter subset for practical deployment.

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights for the 2.4 trillion-parameter model are scheduled for release next week, with the smaller 27B version available sooner for local deployment.

How does Qwen3.8-Max compare to other models like GPT-5.6 or Fable 5?

In benchmark tests, Qwen3.8-Max outperforms Claude models and is close to GPT-5.6 on some tasks, but it trails significantly on deep software engineering benchmarks compared to Fable 5.

What are the main limitations of Qwen3.8-Max?

Its performance on certain deep engineering benchmarks is weaker, and deploying the full 2.4 trillion-parameter model requires substantial hardware infrastructure, limiting immediate practical use.

Source: ThorstenMeyerAI.com

You May Also Like

Rocket Report: Indian startup nears first launch; SpaceX’s millenary milestone

India’s Skyroot Aerospace prepares for its first orbital test flight, while SpaceX achieves a significant milestone with its 1,000th launch, highlighting industry progress.

SpaceX launches Starfall demo mission from Cape Canaveral in Florida

SpaceX successfully launched the Starfall demo mission from Cape Canaveral, Florida, marking a key step in its testing program. Details on mission objectives remain limited.

OpenEuroLLM. The third path.

European consortium OpenEuroLLM faces compute bottlenecks amid ambitious multilingual LLM goals, highlighting limits of pan-European AI pooling.

Évian and the Fallout: What Europe Actually Wants From Amodei, Hassabis, and Altman

European leaders outlined key demands from AI giants Amodei, Hassabis, and Alt after US export controls impact models, emphasizing sovereignty and safety.