Morning Edition · Monday, June 15, 2026Published at 3:00 AM EDT · New York
LFM2.5-8B-A1B activates roughly 1 billion of its 8 billion parameters per token and enables reasoning by default, targeting consumer-device inference.

Liquid AI released LFM2.5-8B-A1B, a mixture-of-experts (MoE) language model with 8 billion total parameters and about 1 billion active parameters per token, according to a post from the AI ML Big Data channel. The stated design goal is on-device inference on laptops and smartphones, with chain-of-thought reasoning enabled by default. The release continues the company's LFM2 line.
The architecture is the point of interest. A sparse mixture-of-experts model with a small active-parameter count keeps per-token computation and memory bandwidth low, which is what matters for local inference, while the larger total parameter pool preserves capacity. An active count near 1 billion places the per-token cost within the range that recent mobile neural processing units and laptop accelerators can sustain, which is the practical threshold for usable local latency.
These are vendor figures from a launch announcement, not independent measurements. The claims to verify are the quantized memory footprint on real consumer hardware, tokens per second under realistic context lengths, and whether the default reasoning mode holds up on standard math and code benchmarks against similarly sized dense models. Until third parties publish numbers, the stated parameter counts describe the design, not delivered quality.
What this means
Small-active mixture-of-experts is becoming the dominant approach for on-device models because it separates capacity from per-token cost. If the reasoning-by-default claim survives independent testing, it moves more agentic workloads off the cloud and onto the device.
What to watch
Observations to monitor, not financial advice.
Start a discussion in Townsquare.
More from this edition
Source: Polylog editors
Comments
0No comments yet.