Morning Edition · Tuesday, August 4, 2026Published at 2:16 AM EDT · New York
Alibaba Ships Qwen3.8-Max at 2.4 Trillion Parameters and Promises the Weights Next Week
The model is priced at $2 and $6 per million input and output tokens, roughly a fifth of Anthropic's top tier, and Alibaba says a Max-class Qwen will be downloadable for the first time.

Alibaba released Qwen3.8-Max on August 3 through Alibaba Cloud's Model Studio. It is a mixture-of-experts (MoE) model with 2.4 trillion total parameters and about 95 billion active per token, a one-million-token context window, and text, image and video input. The company says it will publish the weights on Hugging Face and ModelScope next week, which would make it the first Max-tier Qwen available for download.
The benchmark table comes from Alibaba. On PaperBench, which tests reproduction of research papers, Qwen3.8-Max reports 93.0 against 90.5 for GPT-5.6 Sol, 88.8 for Claude Fable 5 and 80.3 for Opus 4.8. On IFBench it reports 82.8 against 72.7 for the nearest competitor. On Alibaba's run of Terminal Bench 2.1 it reports 86.6, second to GPT-5.6 Sol at 88.8. The longer-horizon claims warrant more caution: a 16-day autonomous coding run and a chip-design flow of more than 500 steps that Alibaba says cut die area by 81 percent. None of that has been reproduced outside the company.
Pricing requires no leaderboard to verify. Application programming interface (API) access runs $2 per million input tokens and $6 per million output tokens, against $10 and $50 for Anthropic's Fable 5. A small independent test posted by the Atomic Chat team compared three interactive three-dimensional physics scenes generated by each model and put Qwen ahead at roughly one-seventh the cost. Alibaba's own launch video leads with chip design and biology, which Russian-language coverage read as a deliberate reply to United States export restrictions on semiconductors.
- If true, who benefits
Alibaba, which converts a downloadable frontier-class model into evidence that United States chip export controls have not held Chinese labs a generation behind, and buyers of agentic coding capacity who gain leverage against Anthropic and OpenAI per-token pricing.
- The nuance
The release, the 2.4-trillion-parameter count and the $2 and $6 pricing are corroborated by multiple independent outlets, but every benchmark figure and the 16-day autonomous run come from Alibaba's own harness, and the promised weight drop is not entirely Alibaba's to make, since Beijing's commerce ministry has held talks with Alibaba, ByteDance and Z.ai about restricting overseas access to flagship models, including open-weight releases.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
If the weights ship as promised, near-frontier agentic coding becomes something a team can download and run rather than a metered call to a vendor in San Francisco, and the competitive pressure falls on closed vendors selling long-horizon agent capability at premium per-token rates. Anthropic and OpenAI hold the top of the coding leaderboards for now, but they lose the argument that only a closed model can sustain a single task for days. Two outcomes are possible. Independent runs of Terminal Bench and SWE-bench confirm the table, in which case Alibaba has cut frontier pricing by roughly five times. Or the gap reappears on evaluations Alibaba did not choose, in which case this is a strong mid-tier model with an aggressive price.
What to watch
- Whether the Hugging Face and ModelScope weights actually appear next week, since a delayed or partial release would indicate the open-weight framing is marketing rather than a plan.
- Independent reruns of Terminal Bench 2.1 and PaperBench by evaluation labs that did not receive the model early, the only way to separate a genuine capability difference from a favorable test harness.
- Whether Anthropic or OpenAI respond with price cuts on their agentic tiers, which would confirm that Chinese open-weight pricing now limits closed-model margins.
Observations to monitor, not financial advice.
Source: Polylog editors
Part of a tracked trend
Open-Weight Models Close the Gap With Closed Frontier Labs
Over the next 3-9 months, open-weight releases with downloadable weights, long context, and strong agentic/coding performance increasingly match closed frontier models on practical work, eroding the closed-lab moat.
More from this edition
- OpenAI Publishes Lean-Checked Proofs for Ten Open Mathematics Problems From an Unreleased Model
- OpenAI Publishes Internal Messages to Rebut Apple's Trade-Secret Suit
- Tencent's Hyra Agent Claims Record Results on 29 of 55 Open Mathematics Problems
- OpenAI Details the Full-Duplex Stack Behind GPT-Live, Built in Six Months
- An Open-Source Runtime Streams Mixture-of-Experts Weights From SSD to Run an 80-Billion-Parameter Qwen on a Mac
- Meta Sells Its Best Model by the Token After Years of Giving Weights Away
- New Papers Push Back on Paying Frontier Prices to Grade Model Output
- Meta's Segmentation and Vision Models Move Into Assistive Robotics and National Laboratory Science
- Researchers Propose an Executable Benchmark for the Decisions Agents Make Before They Answer
- Agent Tooling Turns Toward Traces and Verified Skill Claims
- OpenAI Publishes a Telco Deployment With Revenue Numbers Attached
Comments
0No comments yet.