📊 Full opportunity report: The Numbers Behind Qwen3.8-Max’s AI Capabilities: What’s The Truth? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba has confirmed the specifications of Qwen3.8-Max, including its 2.4 trillion total parameters and 95 billion active parameters. The model’s benchmark results show strong performance, but only on selected tests. Open weights will be released next week, with a smaller, deployable 27B version also coming.
Alibaba has officially confirmed the specifications of its flagship AI model, Qwen3.8-Max, revealing it has 2.4 trillion total parameters and performs strongly on several benchmarks. This marks a significant milestone in large language model development and opens the door for broader access to high-capacity AI models.
On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, previously previewed in July. The model is built on the Qwen3.5 architecture, employs sparse mixture-of-experts, and features a 95 billion active-parameter count, indicating that only about 4% of its total parameters are active per token. The model supports multimodal inputs—text, images, videos—and outputs text.
The benchmark results show the model outperforming several competitors on key tests: it scores 86.6 on Terminal-Bench 2.1, surpassing Claude models and only trailing GPT-5.6. It tops the PaperBench at 93.0 and demonstrates notable strength in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5. However, it trails significantly on deep software engineering benchmarks like SWE-bench Pro, where it scores 67.7 compared to Fable 5’s 80.0, highlighting areas for improvement.
The open weights for the model’s 2.4 trillion parameters are scheduled for release next week, although the current accessible version is a 95 billion-parameter subset optimized for deployment. Alibaba also announced a smaller, 27-billion-parameter version, Qwen3.8-27B, suitable for local deployment and running on a single high-memory machine, with details to be published soon.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba's Model Specifications and Benchmark Results
The confirmation of Qwen3.8-Max’s parameters and benchmark performance underscores Alibaba’s advancing position in large language models, especially in multimodal and agentic tasks. The open release of weights next week could impact AI deployment strategies, enabling wider access to high-capacity models and fostering innovation in AI applications. However, the model’s limitations on deep software engineering benchmarks suggest ongoing challenges in specific technical domains.
For developers and researchers, the availability of a 95 billion active-parameter model with open weights offers new opportunities for customization and integration, though practical deployment remains complex due to hardware requirements. The smaller 27B model aims to bridge that gap, providing a more accessible option for local inference and daily use.

Acer Aspire Go 15 AI Ready Laptop | 15.6" FHD (1920 x 1080) IPS Display | AMD Ryzen 7 7730U | AMD Radeon Graphics | 16GB DDR4 | 512GB PCIe Gen4 SSD | Wi-Fi 6 | Windows 11 Home | AG15-42P-R9FW
- Powerful AMD Ryzen 7 Processor: Smooth, responsive performance for multitasking
- 16GB DDR4 Memory: Enhanced multitasking and faster data access
- 512GB PCIe Gen4 SSD: Fast storage for quick file access
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background and Development Timeline of Alibaba’s AI Models
Alibaba’s AI model journey has been marked by strategic previews and selective disclosures. In July, the company previewed Qwen3.8-Max during the World AI Conference in Shanghai, initially without detailed specifications. The model was identified as “kaleb” on the Code Arena leaderboard, and Alibaba confirmed its identity shortly after. The model’s capabilities and benchmark performance were kept under wraps until today’s full disclosure.
Prior to Qwen3.8-Max, Alibaba released smaller models like Kimi K3 and has steadily increased model size and complexity, emphasizing multimodal and agentic capabilities. The recent release aligns with broader industry trends toward larger, more capable models with open weights, but with a cautious approach to transparency and licensing.
The recent performance claims and benchmark results are part of a broader strategy to establish Alibaba as a leading AI innovator, competing with US and European firms, while managing the risks associated with large-scale model deployment.
"We are committed to open weights and advancing AI accessibility, with the full release scheduled for next week."
— Alibaba spokesperson

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Questions About Model Deployment and Capabilities
It is still unclear how the open weights will be licensed and whether they will be fully open-source or subject to restrictions. The practical implications of deploying a 2.4 trillion-parameter model remain uncertain, given hardware limitations and infrastructure requirements. Additionally, the performance of the smaller 27B version in real-world scenarios is yet to be demonstrated, and whether it can retain the agentic improvements seen in the flagship model is still unknown.
Further details about licensing terms, deployment options, and long-term performance are expected in the coming weeks.

Design Beyond Devices: Creating Multimodal, Cross-Device Experiences
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Alibaba’s AI Model Access and Adoption
Next week, Alibaba plans to release the full 2.4 trillion-parameter model weights, enabling researchers and developers to experiment with the model directly. The smaller Qwen3.8-27B will also become available for local deployment, offering a practical option for daily AI tasks. Industry observers will closely monitor how these models perform in diverse applications and whether licensing terms facilitate widespread adoption.
Further benchmark results and technical documentation are anticipated, providing clarity on the model’s capabilities and limitations, and shaping future AI development strategies.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What are the main capabilities of Qwen3.8-Max?
Qwen3.8-Max features multimodal input support (text, images, videos), strong benchmark performance in tasks like multimodal reasoning and agentic execution, and a 95-billion active-parameter subset for practical deployment.
When will Alibaba release the open weights for Qwen3.8-Max?
The open weights for the 2.4 trillion-parameter model are scheduled for release next week, with the smaller 27B version available sooner for local deployment.
How does Qwen3.8-Max compare to other models like GPT-5.6 or Fable 5?
In benchmark tests, Qwen3.8-Max outperforms Claude models and is close to GPT-5.6 on some tasks, but it trails significantly on deep software engineering benchmarks compared to Fable 5.
What are the main limitations of Qwen3.8-Max?
Its performance on certain deep engineering benchmarks is weaker, and deploying the full 2.4 trillion-parameter model requires substantial hardware infrastructure, limiting immediate practical use.
Source: ThorstenMeyerAI.com