Polylog
The Polylog AI Intelligence Brief

Morning Edition · Saturday, August 1, 2026Published at 1:44 AM EDT · New York

DeepSeek's Smaller V4-Flash Beats Its Own Flagship on Agent Benchmarks at a Fifth of Frontier Cost

The 284-billion-parameter mixture-of-experts model, with 13 billion active parameters, reports 82.7 on Terminal Bench 2.1, above its larger V4-Pro-Preview and just under Anthropic's Opus 4.8.

DeepSeek's Smaller V4-Flash Beats Its Own Flagship on Agent Benchmarks at a Fifth of Frontier Cost

DeepSeek has opened public beta access to a re-trained build of its V4-Flash model, designated V4-Flash-0731, and reports that it outperforms the company's own larger V4-Pro-Preview across all nine agent and coding benchmarks it published. On Terminal Bench 2.1, a test of command-line agentic work, the company puts V4-Flash at 82.7, up from 61.8 for the earlier Flash preview and above the 72.1 it reports for V4-Pro-Preview. On DeepSWE, the re-training moved the score from 7.3 to 54.4, according to TechTimes.

The architecture is unchanged from the Flash preview: a 284-billion-parameter mixture-of-experts model with 13 billion active parameters and a one-million-token context window. What changed is post-training. The gains come from a re-run of the alignment and reinforcement-learning stage rather than a larger base model, which is why a smaller model can exceed a bigger one from the same lab. DeepSeek says the API is directly compatible with the Responses API format and priced at roughly $0.14 per million input tokens on a cache miss and $0.28 per million output tokens, as reported in Russian-language coverage.

The skeptical read matters here. The numbers are vendor-reported and await independent reproduction, and 82.7 on Terminal Bench 2.1 still trails the 85.0 DeepSeek itself attributes to Anthropic's Opus 4.8. The point is not that a Chinese lab has taken the outright agentic-coding lead. It is that a Chinese lab is delivering near-frontier agent performance from a 13-billion-active-parameter model at a fraction of Western closed-model pricing.

Veracity: Plausible
61/100
If true, who benefits

DeepSeek and China's open-weight push, plus buyers of cheap agent inference, at the expense of Western closed-lab pricing power.

The nuance

All nine scores, including the 82.7 on Terminal Bench 2.1, are DeepSeek's own numbers on benchmarks it selected, with no independent reproduction yet.

An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.

What this means

The competitive axis is shifting from raw scale to post-training efficiency and price. A 13-billion-active-parameter model reaching near-frontier agent scores compresses the cost of serving reasoning and coding agents, which pressures the margins of closed labs that charge multiples more per token. The channel of exposure is distribution and unit economics. Teams building coding agents can now run a credible open-stack alternative at DeepSeek's stated prices, and every closed vendor must justify its premium on capability gaps that independent benchmarks have not yet confirmed.

What to watch

  • Independent reproduction of the Terminal Bench 2.1 and DeepSWE numbers on held-out tasks, which would confirm or refute the claim that a smaller model genuinely surpassed the larger one.
  • Whether V4-Flash's one-million-token context holds accuracy on long-horizon agent runs in practice, since context-length claims often degrade on retrieval and multi-step tool use.

Observations to monitor, not financial advice.

2 sources

Synthesized from: Polylog editors · TechTimes

Part of a tracked trend

Chinese Labs Reach Frontier Parity

Chinese labs increasingly match or beat United States frontier offerings on independent benchmarks across modalities, competing on closed metered APIs as well as open weights.