Morning Edition · Wednesday, July 15, 2026Published at 1:32 AM EDT · New York
The Apache-licensed build compresses Qwen3.6-27B from roughly 54 gigabytes to 3.9 and, PrismML says, keeps about 90 percent of full-precision quality.
PrismML on July 14 released Bonsai 27B, 1-bit and 1.58-bit ternary quantizations of Alibaba's Qwen3.6-27B base, and published the weights on Hugging Face under an Apache 2.0 license. The company describes it as the first 27-billion-parameter model that runs locally on a phone.
The compression figures matter most. A 27.8-billion-parameter model at 16-bit precision needs roughly 54 gigabytes of memory before runtime overhead. The 1-bit binary variant measures about 3.9 gigabytes, small enough to fit in the unified memory of current high-end phones and mainstream laptops. PrismML reports roughly 11 tokens per second on an iPhone 17 Pro, with the model keeping its multimodal input, tool use, and reasoning traces, and it ships builds for llama.cpp on CUDA and Metal and for Apple's MLX.
The claim to treat with caution is fidelity. PrismML states the 1-bit build preserves about 90 percent of full-precision quality "across benchmarks," but it has not named the benchmark suite, the baseline scores, or the tasks where the degradation concentrates. Extreme quantization typically hurts long-context recall, multi-step arithmetic, and code the most, and a single top-line percentage does not settle any of that. The weights are downloadable, so independent reproduction on MMLU-Pro, GPQA, or coding suites is the decisive test.
The economic advantage is why engineers should take note regardless. If a 27B-class model runs acceptably on a handset, the marginal cost of many inference calls drops toward zero and leaves the cloud entirely, and the base model is Chinese open weights rather than a US frontier API.
PrismML wins attention and Alibaba gains reach as its open Chinese base becomes the on-device standard, while cloud vendors that bill per token lose the mid-size inference market.
Part of a tracked trend
AI Inference Shifts to Consumer Devices
Over the next 3-6 months, smaller efficient architectures and inference-cost optimizations push capable AI off the cloud and onto laptops, phones, and mobile NPUs.
Start a discussion in Townsquare.
More from this edition
Independent testers reproduce roughly 89.5 percent retention, but PrismML has not named the benchmark suite or the tasks where degradation concentrates, and "first 27B on a phone" is an unverified vendor superlative.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Aggressive post-training quantization is the mechanism moving frontier-adjacent capability off metered cloud APIs and onto devices the user already owns. If the 90 percent retention claim survives independent testing, the parties exposed are cloud inference vendors whose revenue depends on per-token billing of mid-size models, and the parties that gain are device makers, Apple and Qualcomm silicon, and Chinese open-weight suppliers like Alibaba whose base models become the foundation for the on-device tier.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Polylog editors · PrismML · MarkTechPost
Comments
0No comments yet.