The Polylog AI Intelligence Brief

Morning Edition · Tuesday, August 4, 2026Published at 2:16 AM EDT · New York

Alibaba Ships Qwen3.8-Max at 2.4 Trillion Parameters and Promises the Weights Next Week

The model is priced at $2 and $6 per million input and output tokens, roughly a fifth of Anthropic's top tier, and Alibaba says a Max-class Qwen will be downloadable for the first time.

Alibaba Ships Qwen3.8-Max at 2.4 Trillion Parameters and Promises the Weights Next Week

Alibaba released Qwen3.8-Max on August 3 through Alibaba Cloud's Model Studio. It is a mixture-of-experts (MoE) model with 2.4 trillion total parameters and about 95 billion active per token, a one-million-token context window, and text, image and video input. The company says it will publish the weights on Hugging Face and ModelScope next week, which would make it the first Max-tier Qwen available for download.

The benchmark table comes from Alibaba. On PaperBench, which tests reproduction of research papers, Qwen3.8-Max reports 93.0 against 90.5 for GPT-5.6 Sol, 88.8 for Claude Fable 5 and 80.3 for Opus 4.8. On IFBench it reports 82.8 against 72.7 for the nearest competitor. On Alibaba's run of Terminal Bench 2.1 it reports 86.6, second to GPT-5.6 Sol at 88.8. The longer-horizon claims warrant more caution: a 16-day autonomous coding run and a chip-design flow of more than 500 steps that Alibaba says cut die area by 81 percent. None of that has been reproduced outside the company.

Pricing requires no leaderboard to verify. Application programming interface (API) access runs $2 per million input tokens and $6 per million output tokens, against $10 and $50 for Anthropic's Fable 5. A small independent test posted by the Atomic Chat team compared three interactive three-dimensional physics scenes generated by each model and put Qwen ahead at roughly one-seventh the cost. Alibaba's own launch video leads with chip design and biology, which Russian-language coverage read as a deliberate reply to United States export restrictions on semiconductors.

Veracity: Corroborated
79/100
If true, who benefits

Alibaba, which converts a downloadable frontier-class model into evidence that United States chip export controls have not held Chinese labs a generation behind, and buyers of agentic coding capacity who gain leverage against Anthropic and OpenAI per-token pricing.

The nuance

The release, the 2.4-trillion-parameter count and the $2 and $6 pricing are corroborated by multiple independent outlets, but every benchmark figure and the 16-day autonomous run come from Alibaba's own harness, and the promised weight drop is not entirely Alibaba's to make, since Beijing's commerce ministry has held talks with Alibaba, ByteDance and Z.ai about restricting overseas access to flagship models, including open-weight releases.

An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.

What this means

If the weights ship as promised, near-frontier agentic coding becomes something a team can download and run rather than a metered call to a vendor in San Francisco, and the competitive pressure falls on closed vendors selling long-horizon agent capability at premium per-token rates. Anthropic and OpenAI hold the top of the coding leaderboards for now, but they lose the argument that only a closed model can sustain a single task for days. Two outcomes are possible. Independent runs of Terminal Bench and SWE-bench confirm the table, in which case Alibaba has cut frontier pricing by roughly five times. Or the gap reappears on evaluations Alibaba did not choose, in which case this is a strong mid-tier model with an aggressive price.

What to watch

  • Whether the Hugging Face and ModelScope weights actually appear next week, since a delayed or partial release would indicate the open-weight framing is marketing rather than a plan.
  • Independent reruns of Terminal Bench 2.1 and PaperBench by evaluation labs that did not receive the model early, the only way to separate a genuine capability difference from a favorable test harness.
  • Whether Anthropic or OpenAI respond with price cuts on their agentic tiers, which would confirm that Chinese open-weight pricing now limits closed-model margins.

Observations to monitor, not financial advice.

1 source

Source: Polylog editors

Part of a tracked trend

Open-Weight Models Close the Gap With Closed Frontier Labs

Over the next 3-9 months, open-weight releases with downloadable weights, long context, and strong agentic/coding performance increasingly match closed frontier models on practical work, eroding the closed-lab moat.

Share this article

Comments

0

No comments yet.