Morning Edition · Saturday, September 12, 2026Published at 2:23 AM EDT · New York
Fugu Max prices at $2 per million input tokens and $6 per million output tokens by routing work to the smallest model capable of handling it, including NVIDIA Nemotron open-weight models.

Sakana AI has released Fugu Max and Fugu Ultra v2, two configurations of the same orchestration architecture built for different priorities. Fugu is not a foundation model. It is a learned router that breaks a request into pieces, assigns e…
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
The Global Intelligence Brief stays free.
Part of a tracked trend
The Inference-Cost Efficiency Race
Techniques that cut tokens generated and KV-cache memory per query will keep compressing the marginal cost of serving reasoning models, making inference efficiency a recurring competitive axis alongside raw capability.
Start a discussion in Townsquare.
More from this edition
Comments
0No comments yet.