Knowledge Base

Bonsai 27B Release (PrismML)

Bonsai 27B Release (PrismML)

Source: https://x.com/i/status/2077087413076176953

📌 Xenova congratulated PrismML on the July 14, 2026 release of Bonsai 27B, a compressed 27B-class multimodal LLM derived from Qwen3.6-27B that fits on phones and laptops. The tweet points to a Hugging Face model collection and a browser WebGPU demo for running the 1-bit variant locally.

📦 Two Size Tiers

Ternary Bonsai 27B (~5.9 GB, 1.71 bits/weight) targets laptop quality; 1-bit Bonsai 27B (~3.9 GB, 1.125 bits/weight) is the phone-oriented footprint variant.

📱 Phone-First Milestone

At ~4 GB, the 1-bit build is the first 27B-class model to run on-device on an iPhone 17 Pro, using MLX Swift at roughly 11 tok/s.

🧠 Intelligence Retained

Across 15 thinking-mode benchmarks, Ternary retains ~95% and 1-bit ~90% of the full-precision Qwen3.6-27B baseline, with math and coding hit least.

🤖 Multimodal & Agentic

Supports 262K-token context, vision (4-bit tower), tool calling, and multi-step agentic workflows—aimed at local, private, offline assistants.

🌐 Try It Now

Weights ship on Hugging Face (GGUF, MLX, AWQ); Xenova's linked WebGPU Space runs the 1-bit model in-browser via custom kernels.

Key facts

Fact Value
Announced July 14, 2026
Base Model Qwen3.6-27B (~27.3B parameters)
1-bit Footprint ~3.9 GB (vs ~54 GB FP16)
Ternary Footprint ~5.9 GB
License Apache 2.0
Platforms Apple MLX, NVIDIA CUDA/llama.cpp, browser WebGPU

Details

PrismML's Bonsai 27B applies extreme low-bit weight compression—binary {−1, +1} or ternary {−1, 0, +1} with FP16 group-wise scaling—end-to-end across embeddings, attention, MLPs, and the LM head. Unlike post-hoc quantization, the representation is native: weights stay packed at inference with custom kernels in PrismML forks of llama.cpp, MLX, and mlx-swift. The vision tower ships separately in compact 4-bit HQQ and loads only when images are provided.

The release is framed around "intelligence density": capability per gigabyte deployed. On an RTX 5090, 1-bit reaches up to ~163 tok/s and Ternary ~134 tok/s; on M5 Max, ~87 and ~58 tok/s respectively. A DSpark speculative-decoding drafter adds a lossless ~1.37× CUDA decode speedup. Benchmarks span knowledge, math, coding, instruction following, tool calling (BFCL v3, τ²-Bench), and vision (MMMU-Pro, OCRBench).

Xenova's post amplifies community access: the Hugging Face collection bundles GGUF, MLX, AWQ, and unpacked weights plus the webml-community WebGPU demo for in-browser inference. Xenova (Transformers.js maintainer) also submitted the story to Hacker News, where it drew heavy discussion—including reports that Apple is in talks with PrismML about on-device compression for iPhone.

PrismML emerged from Caltech researchers, backed by Khosla Ventures, Cerberus, Google, and Samsung. The company positions Bonsai 27B for hybrid cloud/local agentic systems: route privacy-sensitive and high-volume steps to a capable on-device model and reserve frontier cloud models for the hardest tasks, with zero marginal per-token cost for long multi-step loops.

Sources