# Alibaba Ships Qwen3.8-Max at 2.4 Trillion Parameters and Promises the Weights Next Week

The model is priced at $2 and $6 per million input and output tokens, roughly a fifth of Anthropic's top tier, and Alibaba says a Max-class Qwen will be downloadable for the first time.

- Published: 2026-08-04T06:16:46.804Z
- Canonical: https://polylog.news/ai/2026-08-04/alibaba-ships-qwen3-8-max-at-2-4-trillion-parameters-and-pro
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news)

Alibaba released [Qwen3.8-Max](https://the-decoder.com/alibabas-open-weight-qwen3-8-max-takes-on-long-horizon-ai-tasks-with-2-4-trillion-parameters/) on August 3 through Alibaba Cloud's Model Studio. It is a mixture-of-experts (MoE) model with 2.4 trillion total parameters and about 95 billion active per token, a one-million-token context window, and text, image and video input. The company says it will publish the weights on Hugging Face and ModelScope [next week](https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/), which would make it the first Max-tier Qwen available for download.

The benchmark table comes from Alibaba. On [PaperBench](https://apidog.com/blog/qwen-3-8-benchmarks/), which tests reproduction of research papers, Qwen3.8-Max reports 93.0 against 90.5 for GPT-5.6 Sol, 88.8 for Claude Fable 5 and 80.3 for Opus 4.8. On IFBench it reports 82.8 against 72.7 for the nearest competitor. On Alibaba's run of Terminal Bench 2.1 it reports 86.6, second to GPT-5.6 Sol at 88.8. The longer-horizon claims warrant more caution: a [16-day autonomous coding run](https://www.developer-tech.com/news/alibaba-qwen3-8-max-claims-16-day-autonomous-coding-run/) and a chip-design flow of more than 500 steps that Alibaba says cut die area by 81 percent. None of that has been reproduced outside the company.

Pricing requires no leaderboard to verify. [Application programming interface (API) access runs $2 per million input tokens and $6 per million output tokens](https://www.qubrid.com/blog/qwen-38-max-api-is-now-live-on-qubrid-ai-benchmarks-pricing-and-how-to-actually-run-it), against $10 and $50 for Anthropic's Fable 5. A small independent test posted by the Atomic Chat team compared three interactive three-dimensional physics scenes generated by each model and [put Qwen ahead at roughly one-seventh the cost](https://t.me/ai_machinelearning_big_data/10643). Alibaba's own launch video leads with chip design and biology, [which Russian-language coverage read as a deliberate reply](https://t.me/ai_machinelearning_big_data/10638) to United States export restrictions on semiconductors.

## What this means

If the weights ship as promised, near-frontier agentic coding becomes something a team can download and run rather than a metered call to a vendor in San Francisco, and the competitive pressure falls on closed vendors selling long-horizon agent capability at premium per-token rates. Anthropic and OpenAI hold the top of the coding leaderboards for now, but they lose the argument that only a closed model can sustain a single task for days. Two outcomes are possible. Independent runs of Terminal Bench and SWE-bench confirm the table, in which case Alibaba has cut frontier pricing by roughly five times. Or the gap reappears on evaluations Alibaba did not choose, in which case this is a strong mid-tier model with an aggressive price.

## What to watch

- Whether the Hugging Face and ModelScope weights actually appear next week, since a delayed or partial release would indicate the open-weight framing is marketing rather than a plan.
- Independent reruns of Terminal Bench 2.1 and PaperBench by evaluation labs that did not receive the model early, the only way to separate a genuine capability difference from a favorable test harness.
- Whether Anthropic or OpenAI respond with price cuts on their agentic tiers, which would confirm that Chinese open-weight pricing now limits closed-model margins.
