Bonsai 27B Release (PrismML)
Bonsai 27B Release (PrismML)
Source: https://x.com/i/status/2077087413076176953
📌 Xenova congratulated PrismML on the July 14, 2026 release of Bonsai 27B, a compressed 27B-class multimodal LLM derived from Qwen3.6-27B that fits on phones and laptops. The tweet points to a Hugging Face model collection and a browser WebGPU demo for running the 1-bit variant locally.
📦 Two Size Tiers
Ternary Bonsai 27B (~5.9 GB, 1.71 bits/weight) targets laptop quality; 1-bit Bonsai 27B (~3.9 GB, 1.125 bits/weight) is the phone-oriented footprint variant.
📱 Phone-First Milestone
At ~4 GB, the 1-bit build is the first 27B-class model to run on-device on an iPhone 17 Pro, using MLX Swift at roughly 11 tok/s.
🧠 Intelligence Retained
Across 15 thinking-mode benchmarks, Ternary retains ~95% and 1-bit ~90% of the full-precision Qwen3.6-27B baseline, with math and coding hit least.
🤖 Multimodal & Agentic
Supports 262K-token context, vision (4-bit tower), tool calling, and multi-step agentic workflows—aimed at local, private, offline assistants.
🌐 Try It Now
Weights ship on Hugging Face (GGUF, MLX, AWQ); Xenova's linked WebGPU Space runs the 1-bit model in-browser via custom kernels.
Key facts
| Fact | Value |
|---|---|
| Announced | July 14, 2026 |
| Base Model | Qwen3.6-27B (~27.3B parameters) |
| 1-bit Footprint | ~3.9 GB (vs ~54 GB FP16) |
| Ternary Footprint | ~5.9 GB |
| License | Apache 2.0 |
| Platforms | Apple MLX, NVIDIA CUDA/llama.cpp, browser WebGPU |
Details
PrismML's Bonsai 27B applies extreme low-bit weight compression—binary {−1, +1} or ternary {−1, 0, +1} with FP16 group-wise scaling—end-to-end across embeddings, attention, MLPs, and the LM head. Unlike post-hoc quantization, the representation is native: weights stay packed at inference with custom kernels in PrismML forks of llama.cpp, MLX, and mlx-swift. The vision tower ships separately in compact 4-bit HQQ and loads only when images are provided.
The release is framed around "intelligence density": capability per gigabyte deployed. On an RTX 5090, 1-bit reaches up to ~163 tok/s and Ternary ~134 tok/s; on M5 Max, ~87 and ~58 tok/s respectively. A DSpark speculative-decoding drafter adds a lossless ~1.37× CUDA decode speedup. Benchmarks span knowledge, math, coding, instruction following, tool calling (BFCL v3, τ²-Bench), and vision (MMMU-Pro, OCRBench).
Xenova's post amplifies community access: the Hugging Face collection bundles GGUF, MLX, AWQ, and unpacked weights plus the webml-community WebGPU demo for in-browser inference. Xenova (Transformers.js maintainer) also submitted the story to Hacker News, where it drew heavy discussion—including reports that Apple is in talks with PrismML about on-device compression for iPhone.
PrismML emerged from Caltech researchers, backed by Khosla Ventures, Cerberus, Google, and Samsung. The company positions Bonsai 27B for hybrid cloud/local agentic systems: route privacy-sensitive and high-volume steps to a capable on-device model and reserve frontier cloud models for the hardest tasks, with zero marginal per-token cost for long multi-step loops.