Morning Edition · Thursday, August 13, 2026Published at 2:26 AM EDT · New York
DeepSeek V4 Pro 0813 posted 87.9 on Terminal-Bench 2.1 at $0.87 per million output tokens, Grok 4.6 took the top spot on a knowledge-work evaluation, and Alibaba moved its 2.4-trillion-parameter model to open weights.

Three significant model releases landed within 24 hours. DeepSeek moved V4 Pro 0813 to general availability, SpaceXAI released Grok 4.6, and Alibaba's Qwen team committed to releasing downloadable weights for its largest model. The Russian-language channel AI ML Big Data summarized all three together, and the pattern across them is consistent: the capability gap at the top is now measured in fractions of a benchmark point, while the price gap is measured in multiples.
DeepSeek V4 Pro 0813 is listed on OpenRouter with a 1,048,576-token context window and a maximum output of 384,000 tokens, priced at $0.435 per million input tokens and $0.87 per million output tokens. On Terminal-Bench 2.1 the model scores 87.9, 0.1 point behind Fable 5 at 88.0 and above Opus 4.8 at 85.0. The larger gain is against DeepSeek's own April preview version: its DeepSWE score rose from 12.8 to 62.7 and its CyberGym score from 52.7 to 83.3. Those figures come from DeepSeek and have not been independently verified.
Grok 4.6 arrived the same day from SpaceXAI, the unit that absorbed xAI after SpaceX acquired it in an all-stock deal. It carries a 500,000-token context window at $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.5. On the Artificial Analysis Intelligence Index it scored 61, tying GPT-5.6 Sol Max and sitting one point behind Fable 5 Max. It leads on the GDPVal-AA v2 evaluation, scoring 1,753 against Fable 5 Max's 1,741, but trails on CursorBench, DeepSWE, FrontierCode, APEX-Agents, Terminal-Bench and APEX-SWE. Taken together, the results describe a strong knowledge-work model, not a coding leader.
Alibaba's Qwen3.8-Max is a 2.4-trillion-parameter sparse mixture-of-experts model with roughly 95 billion parameters active per token, 512 experts, and a one-million-token context window. It scores 86.6 on Terminal-Bench 2.1 against GPT-5.6 Sol's 88.8. Alibaba announced the architecture on August 3 and committed to publishing weights for both the full Max model and a smaller 27-billion-parameter version on Hugging Face and ModelScope. The license has not been disclosed. Qwen 3.5 and 3.6 shipped under the permissive Apache 2.0 license, which is a precedent, not a guarantee. A 2.4-trillion-parameter model is downloadable in name only, since very few organizations have the hardware to run it, so the 27-billion-parameter release is the one that will actually change what most teams can use.
Part of a tracked trend
Open-Weight Models Close the Gap With Closed Frontier Labs
Over the next 3-9 months, open-weight releases with downloadable weights, long context, and strong agentic/coding performance increasingly match closed frontier models on practical work, eroding the closed-lab moat.
Start a discussion in Townsquare.
More from this edition
DeepSeek and Alibaba, whose price and open-weights positioning erodes the pricing power of closed mid-tier models, and the inference providers and enterprise buyers who gain leverage in procurement, while the labs releasing the numbers are the same parties whose valuations depend on them.
Grok 4.6's score of 61 comes from the independent evaluator Artificial Analysis, which also records Grok at 88.4 on Terminal-Bench v2.1 and therefore does not support the claim that it trails there, DeepSeek's agentic gains remain vendor-reported and unreproduced outside the lab with comparators varying across write-ups, and Alibaba's Qwen3.8 weights were announced but not published as of early August with no license disclosed, so "moved to open weights" describes a commitment rather than a completed release.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Within about two points of each other on the leading agentic coding benchmarks, buyers can now choose between models priced at $6 and $0.87 per million output tokens. That narrows how much a closed lab can charge for anything short of frontier-level performance, and it shifts competition toward reliability, tool-use behavior over long tasks, and enterprise controls rather than raw benchmark scores. Chinese labs gain distribution through lower cost, and inference providers gain volume. The vendors most exposed are those selling mid-tier closed models whose main selling point was a benchmark lead now measured in tenths of a point.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Polylog editors · OpenRouter (via Hacker News)
Comments
1Aug 14, 4:00 AM · edited
Alibaba's commitment to releasing 2.4T parameter weights places that capability tier beyond future access policy or export control reach once the download completes.