# DeepSeek V4-Flash Undercuts Western Frontier Models on Cost, With a Token-Verbosity Caveat

The 284-billion-parameter mixture-of-experts model prices at $0.14 and $0.28 per million input and output tokens, but consumed roughly 3.5 times the median tokens to finish a benchmark.

- Published: 2026-08-02T05:44:51.872Z
- Canonical: https://polylog.news/ai/2026-08-02/deepseek-v4-flash-undercuts-western-frontier-models-on-cost
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news), [Artificial Analysis](https://artificialanalysis.ai/models/deepseek-v4-flash)

Early analyses circulating among practitioners [claim DeepSeek's V4-Flash](https://t.me/aipost/7704) completes some benchmark task suites at a total cost roughly 105 times lower than the leading Western frontier model, a figure that depends on both per-token price and how many tokens a model uses to finish a task.

The per-token side is not in dispute. Independent tracking [lists V4-Flash](https://artificialanalysis.ai/models/deepseek-v4-flash) at $0.14 per million input tokens and $0.28 per million output tokens, well under the market median of $0.58 and $2.20, and scores it around 50 on the Artificial Analysis Intelligence Index. It is an efficiency-tuned mixture-of-experts design with 284 billion total and roughly 13 billion active parameters, and it scores 82.7 on Terminal Bench 2.1 for agentic workloads, above DeepSeek's own V4-Pro preview.

The caveat is verbosity. The same tracking notes V4-Flash used about 210 million output tokens to complete the Intelligence Index, roughly 3.5 times the median, so an advertised per-token discount shrinks once the number of tokens a reasoning model uses is counted. The accurate measure is total task cost rather than per-token price, and on cost per completed task the model is cheap but not cheap by the full multiple its per-token rate implies.

## What this means

The channel is inference cost. An open-weight model that ranks below the frontier but is priced this far below Western application programming interfaces (APIs) pressures margins for every closed vendor that sells reasoning tokens, because buyers now compare models on cost per completed task, where an efficient mixture-of-experts with few active parameters comes out ahead even after its higher token use is counted. The exposed parties are metered API sellers whose pricing assumes capability is scarce. The beneficiary is any deployer whose workload tolerates a small quality gap for a large cost cut.

## What to watch

- Whether independent evaluators reproduce the total-cost-per-task gap after normalizing for token consumption, which separates a real efficiency win from a per-token headline.
- Repricing responses from closed vendors, which would signal that price competition is moving from open weights into the metered frontier tier.
