# Three Frontier Models Shipped in One Day, and Two of Them Undercut the Leaders on Price

DeepSeek V4 Pro 0813 posted 87.9 on Terminal-Bench 2.1 at $0.87 per million output tokens, Grok 4.6 took the top spot on a knowledge-work evaluation, and Alibaba moved its 2.4-trillion-parameter model to open weights.

- Published: 2026-08-13T06:26:09.114Z
- Canonical: https://polylog.news/ai/2026-08-13/three-frontier-models-shipped-in-one-day-and-two-of-them-und
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news), [OpenRouter (via Hacker News)](https://openrouter.ai/deepseek/deepseek-v4-pro-0813)

Three significant model releases landed within 24 hours. DeepSeek moved V4 Pro 0813 to general availability, SpaceXAI released Grok 4.6, and Alibaba's Qwen team committed to releasing downloadable weights for its largest model. The Russian-language channel AI ML Big Data [summarized all three together](https://t.me/ai_machinelearning_big_data/10695), and the pattern across them is consistent: the capability gap at the top is now measured in fractions of a benchmark point, while the price gap is measured in multiples.

DeepSeek V4 Pro 0813 is listed on [OpenRouter](https://openrouter.ai/deepseek/deepseek-v4-pro-0813) with a 1,048,576-token context window and a maximum output of 384,000 tokens, priced at $0.435 per million input tokens and $0.87 per million output tokens. On Terminal-Bench 2.1 the model [scores 87.9](https://explainx.ai/blog/deepseek-v4-pro-0813-terminal-bench-cline-august-2026), 0.1 point behind Fable 5 at 88.0 and above Opus 4.8 at 85.0. The larger gain is against DeepSeek's own April preview version: its DeepSWE score rose from 12.8 to 62.7 and its CyberGym score from 52.7 to 83.3. Those figures come from DeepSeek and have not been independently verified.

Grok 4.6 arrived the same day from SpaceXAI, the unit that absorbed xAI after [SpaceX acquired it in an all-stock deal](https://en.wikipedia.org/wiki/SpaceXAI). It carries a 500,000-token context window at $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.5. On the Artificial Analysis Intelligence Index it [scored 61, tying GPT-5.6 Sol Max](https://venturebeat.com/technology/spacexai-debuts-grok-4-6-overtaking-kimi-k3s-performance-and-matching-gpt-5-6-sol-for-worlds-third-best-on-artificial-analysis) and sitting one point behind Fable 5 Max. It leads on the GDPVal-AA v2 evaluation, scoring 1,753 against Fable 5 Max's 1,741, but trails on CursorBench, DeepSWE, FrontierCode, APEX-Agents, Terminal-Bench and APEX-SWE. Taken together, the results describe a strong knowledge-work model, not a coding leader.

Alibaba's Qwen3.8-Max is a 2.4-trillion-parameter sparse mixture-of-experts model with roughly 95 billion parameters active per token, 512 experts, and a one-million-token context window. It scores [86.6 on Terminal-Bench 2.1 against GPT-5.6 Sol's 88.8](https://www.datacamp.com/blog/qwen3-8-max). Alibaba announced the architecture on August 3 and [committed to publishing weights](https://www.scmp.com/tech/article/3362738/alibabas-ai-model-qwen38-max-made-widely-accessible-ahead-open-weights-release) for both the full Max model and a smaller 27-billion-parameter version on Hugging Face and ModelScope. The license has not been disclosed. Qwen 3.5 and 3.6 shipped under the permissive Apache 2.0 license, which is a precedent, not a guarantee. A 2.4-trillion-parameter model is downloadable in name only, since very few organizations have the hardware to run it, so the 27-billion-parameter release is the one that will actually change what most teams can use.

## What this means

Within about two points of each other on the leading agentic coding benchmarks, buyers can now choose between models priced at $6 and $0.87 per million output tokens. That narrows how much a closed lab can charge for anything short of frontier-level performance, and it shifts competition toward reliability, tool-use behavior over long tasks, and enterprise controls rather than raw benchmark scores. Chinese labs gain distribution through lower cost, and inference providers gain volume. The vendors most exposed are those selling mid-tier closed models whose main selling point was a benchmark lead now measured in tenths of a point.

## What to watch

- Whether independent evaluators reproduce DeepSeek's Terminal-Bench and CyberGym scores, since every figure so far comes from the lab that benefits from them.
- The license Alibaba attaches to the Qwen3.8 weights, since an Apache 2.0 release and a restricted enterprise license would lead to very different adoption in regulated industries.
- Whether DeepSeek follows through on its signalled API price increase, which would show current pricing is a push for market share rather than a sustainable cost structure.
