Morning Edition · Tuesday, August 11, 2026Published at 2:26 AM EDT · New York
Salvatore Sanfilippo has written a native Metal inference engine for MiniMax-H3, while a separate MLX port needs about 115 gigabytes of weights and takes just under 45 minutes per clip on an M5 Max.

MiniMax-H3 is a 33-billion-parameter joint video and audio diffusion model, and two independent efforts have now brought it to Apple Silicon without using PyTorch. Salvatore Sanfilippo, the creator of Redis, is building h3.c, a native infer…
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
The Global Intelligence Brief stays free.
Part of a tracked trend
AI Inference Shifts to Consumer Devices
Over the next 3-6 months, smaller efficient architectures and inference-cost optimizations push capable AI off the cloud and onto laptops, phones, and mobile NPUs.
Start a discussion in Townsquare.
More from this edition
Comments
0No comments yet.