Polylog
The Polylog AI Intelligence Brief

Morning Edition · Sunday, August 2, 2026Published at 1:44 AM EDT · New York

DeepSeek V4-Flash Undercuts Western Frontier Models on Cost, With a Token-Verbosity Caveat

The 284-billion-parameter mixture-of-experts model prices at $0.14 and $0.28 per million input and output tokens, but consumed roughly 3.5 times the median tokens to finish a benchmark.

DeepSeek V4-Flash Undercuts Western Frontier Models on Cost, With a Token-Verbosity Caveat

Early analyses circulating among practitioners claim DeepSeek's V4-Flash completes some benchmark task suites at a total cost roughly 105 times lower than the leading Western frontier model, a figure that depends on both per-token price and how many tokens a model uses to finish a task.

The per-token side is not in dispute. Independent tracking lists V4-Flash at $0.14 per million input tokens and $0.28 per million output tokens, well under the market median of $0.58 and $2.20, and scores it around 50 on the Artificial Analysis Intelligence Index. It is an efficiency-tuned mixture-of-experts design with 284 billion total and roughly 13 billion active parameters, and it scores 82.7 on Terminal Bench 2.1 for agentic workloads, above DeepSeek's own V4-Pro preview.

The caveat is verbosity. The same tracking notes V4-Flash used about 210 million output tokens to complete the Intelligence Index, roughly 3.5 times the median, so an advertised per-token discount shrinks once the number of tokens a reasoning model uses is counted. The accurate measure is total task cost rather than per-token price, and on cost per completed task the model is cheap but not cheap by the full multiple its per-token rate implies.

Veracity: Plausible
55/100
If true, who benefits

DeepSeek and the broader argument that Chinese open-weight models reach near-frontier quality far below Western prices, which pressures the pricing power and margins of metered US API vendors.

The nuance

The load-bearing 105-times figure is a per-token, benchmark-specific selection, while independent tracking on Artificial Analysis puts the cost-per-task advantage near 60 percent cheaper than a leading Western model once the model's roughly 3.5-times-higher token use is counted.

An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.

What this means

The channel is inference cost. An open-weight model that ranks below the frontier but is priced this far below Western application programming interfaces (APIs) pressures margins for every closed vendor that sells reasoning tokens, because buyers now compare models on cost per completed task, where an efficient mixture-of-experts with few active parameters comes out ahead even after its higher token use is counted. The exposed parties are metered API sellers whose pricing assumes capability is scarce. The beneficiary is any deployer whose workload tolerates a small quality gap for a large cost cut.

What to watch

  • Whether independent evaluators reproduce the total-cost-per-task gap after normalizing for token consumption, which separates a real efficiency win from a per-token headline.
  • Repricing responses from closed vendors, which would signal that price competition is moving from open weights into the metered frontier tier.

Observations to monitor, not financial advice.

2 sources

Synthesized from: Polylog editors · Artificial Analysis

Part of a tracked trend

The Inference-Cost Efficiency Race

Techniques that cut tokens generated and KV-cache memory per query will keep compressing the marginal cost of serving reasoning models, making inference efficiency a recurring competitive axis alongside raw capability.