PrismML claims first 27B model on iPhone
PrismML's Bonsai 27B, a compressed Qwen3.6 27B, is Apache 2.0 licensed and, the company says, the first 27B-class model to run on an iPhone.
PrismML released Bonsai 27B on July 14, a compressed version of Alibaba’s Qwen3.6 27B that the Caltech spinout says is the first model of that parameter class to run on a phone, fitting the memory budget of the 12GB iPhone 17 Pro and Pro Max. PrismML is led by Caltech professor Babak Hassibi and backed by a $16.25 million seed round from Khosla Ventures and Cerberus, with compute support from Google and Caltech.
The release ships two quantized variants. A “ternary” version, using three-value weights at 5.9GB, is what PrismML calls its quality-oriented option and says retains about 95% of the full-precision model’s performance. A 1-bit version, at 3.9GB, is the smaller, footprint-oriented option, which PrismML says retains about 90%. Both are multimodal, with a compact 4-bit vision component. The weights are released under Apache 2.0 on Hugging Face and GitHub.
Why it matters
If PrismML’s numbers hold up under independent testing, running a 27B-class model entirely on a phone, with no server round trip, is a meaningful jump for on-device AI. For now the 95% and 90% retention figures are PrismML’s own benchmarks, not independently verified, so treat them as a claim rather than a settled result. It is the same on-device push that led Google to put Gemma 4 E2B natively on Pixel 10’s Tensor chip this week, and the same caution that applies to Liquid AI’s own benchmark claims for its 230M model.