The Polylog AI Intelligence Brief

Morning Edition · Tuesday, August 4, 2026Published at 2:16 AM EDT · New York

An Open-Source Runtime Streams Mixture-of-Experts Weights From SSD to Run an 80-Billion-Parameter Qwen on a Mac

Swiftlet keeps only the dense core in memory and fetches each routed expert with a single disk read, putting a 35-billion-parameter model on an iPhone at roughly one token per second.

An Open-Source Runtime Streams Mixture-of-Experts Weights From SSD to Run an 80-Billion-Parameter Qwen on a Mac

A Swift and Metal runtime called Swiftlet applies a straightforward observation about sparse models to deployment. In a mixture-of-experts (MoE) model, only a small fraction of parameters activate per token, so the routed experts do not nee…

Continue the AI Intelligence Brief

Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.

  • 5 AI intelligence signals a day
  • Frontier labs, compute, and chips
  • Model releases and AI infrastructure
  • Source-grounded analysis with confidence labels

The Global Intelligence Brief stays free.

Subscribe for $19/mo

Part of a tracked trend

AI Inference Shifts to Consumer Devices

Over the next 3-6 months, smaller efficient architectures and inference-cost optimizations push capable AI off the cloud and onto laptops, phones, and mobile NPUs.

Share this article

Comments

0

No comments yet.