# Alibaba Ships Qwen3.8-Max, a 2.4-Trillion-Parameter Model It Promises to Open-Weight Next Week

The model activates 95 billion parameters per request and prices at two dollars per million input tokens. It still trails Anthropic's Fable 5 on the hardest public software-engineering benchmark.

- Published: 2026-08-04T06:03:13.954Z
- Canonical: https://polylog.news/ai/2026-08-04/alibaba-ships-qwen3-8-max-a-2-4-trillion-parameter-model-it
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news), [CNBC](https://www.cnbc.com/2026/08/03/alibaba-ai-model-qwen-rival-anthropic.html), [MarkTechPost](https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/)

Alibaba released Qwen3.8-Max on Monday, a mixture-of-experts model with 2.4 trillion total parameters, roughly 95 billion of which activate on any single request. It handles text, images and video, and it supports a context window of up to one million tokens. Alibaba serves it through QwenCloud at two dollars per million input tokens and six dollars per million output tokens, with cached input at twenty-five cents. The company said the weights will be published on Hugging Face and ModelScope next week under an Apache 2.0 license, alongside a much smaller Qwen3.8-27B, which [MarkTechPost reports would make it the first downloadable Max-class Qwen model](https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/).

The company's own numbers put the model ahead of OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 on several coding, multimodal and engineering evaluations, and behind them on parts of general reasoning. The actual difference is narrower than the marketing suggests. Qwen3.8-Max reports 67.7 on SWE-bench Pro, while Anthropic [claims 80.3 on the same benchmark](https://www.anthropic.com/news/claude-fable-5-mythos-5) for Fable 5 and Mythos 5. Every one of these figures comes from a vendor rather than from an independent reproduction, and the weights that would allow outside verification have not been released.

The independent signal so far is informal. A team at Atomic Chat gave Qwen3.8-Max and Fable 5 the same task: three self-contained interactive HTML scenes with three-dimensional physics, including marbles riding a wheel through loops. The team [reported that the Qwen output was better and cost close to seven times less](https://t.me/ai_machinelearning_big_data/10643). That is one prompt run by one team, which is evidence about price and performance rather than proof of capability parity.

The presentation around the release is as deliberate as the model itself. Alibaba's promotional video shows the model spending hours on chip design work before moving to biology tasks, a [pairing that Russian-language AI coverage read as a direct answer to United States export controls](https://t.me/ai_machinelearning_big_data/10638) on semiconductors. Investors responded to the pricing rather than to the benchmark results. Alibaba's Hong Kong-listed shares [closed about 7% higher at HK$125.20](https://invezz.com/news/2026/08/03/alibaba-shares-jump-6-after-launch-of-qwen3-8-max-its-biggest-ai-model-yet/), and its American depositary receipts rose in premarket trading.

## What this means

A downloadable Max-class model at two dollars per million input tokens competes with the closed labs at their weakest point, which is the price of the second-best answer. Enterprises that route bulk agentic and coding traffic through an application programming interface (API) now have a self-hostable option at a fraction of frontier pricing. That pressures per-token margins at OpenAI and Anthropic and strengthens Alibaba's cloud business, where the model works as a customer-acquisition mechanism rather than as the product. The unresolved question is whether the released weights reproduce the claimed scores. If independent runs come in near 67.7 on SWE-bench Pro, Alibaba has a strong value model that still trails the frontier on hard software work. If they come in well below, the vendor numbers were tuned and price is the only real advantage.

## What to watch

- Whether the promised Hugging Face and ModelScope weights actually appear next week, and whether outside groups reproduce the coding scores. A gap between the vendor table and independent runs would be the clearest evidence yet that Max-class benchmark claims need discounting.
- How OpenAI and Anthropic respond on price rather than on capability. A repricing of their mid-tier coding endpoints would confirm that open Chinese weights now set the lowest price enterprises expect to pay per token.
- Whether Western enterprises and governments allow a Chinese-origin open-weight model into production stacks, since self-hosting removes the data-residency objection but not the procurement politics.
