The Polylog AI Intelligence Brief

Morning Edition · Tuesday, August 4, 2026Published at 2:03 AM EDT · New York

An Open-Source Runtime Puts an 80-Billion-Parameter Qwen Model on a Mac in 4.3 Gigabytes of Memory

Swiftlet keeps only the dense core resident and streams routed experts from storage. It runs a 35-billion-parameter model on an iPhone 17 at roughly one token per second.

An Open-Source Runtime Puts an 80-Billion-Parameter Qwen Model on a Mac in 4.3 Gigabytes of Memory

A Swift and Metal runtime called Swiftlet, published this week, targets the mixture-of-experts models in the Qwen3-Next and Qwen3.5/3.6 families and exploits a structural fact about them. Only a small fraction of the weights participate in…

Continue the AI Intelligence Brief

Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.

  • 5 AI intelligence signals a day
  • Frontier labs, compute, and chips
  • Model releases and AI infrastructure
  • Source-grounded analysis with confidence labels

The Global Intelligence Brief stays free.

Subscribe for $19/mo

Part of a tracked trend

AI Inference Shifts to Consumer Devices

Over the next 3-6 months, smaller efficient architectures and inference-cost optimizations push capable AI off the cloud and onto laptops, phones, and mobile NPUs.

Share this article

Comments

0

No comments yet.