Knowledge Base

Bonsai 27B: PrismML’s Phone-Sized 27B Model

Bonsai 27B: PrismML’s Phone-Sized 27B Model

Source: https://x.com/i/status/2077710397801386480

📌 Bonsai 27B is PrismML’s low-bit multimodal build of Qwen3.6-27B, compressed to ternary (~5.9 GB) or 1-bit (~3.9 GB) while retaining roughly 90–95% of full-precision scores. The 1-bit variant is pitched as the first 27B-class model to run fully on a consumer smartphone under Apache 2.0.

📦 Extreme compression

Full 16-bit 27B weights are ~54 GB; conventional 4-bit still ~18 GB. PrismML packs weights to ternary {−1,0,+1} or binary {−1,+1}, shrinking the model to ~5.9 GB (1.71 bpw) or ~3.9 GB (1.125 bpw) without changing the Qwen3.6-27B architecture.

📱 Phone-first footprint

iOS apps often sit under ~6 GB RAM, so ordinary 27B builds never fit. The 1-bit Bonsai variant is promoted as the first 27B-class model on a consumer phone—~11 tok/s on iPhone 17 Pro Max via MLX Swift, with Atomic Chat demos on iPhone and Android.

🧠 Workstation-class local use

Supports multi-step reasoning, tool use, long context (~262K), multimodal vision, and agent-style loops fully on-device—no cloud required. Ternary targets laptop-class quality; 1-bit targets phone footprint.

🔓 Open weights & runtimes

Apache 2.0 on Hugging Face (GGUF, MLX, and related formats) with specialized llama.cpp/MLX runtimes for the unusual bit packing. Available via Atomic Chat (iOS/Android) and other local apps.

📊 Quality retained

Publisher-reported retention: ternary ~94.6% of FP16 benchmarks; 1-bit ~89.5%. Together they push private offline AI from small mobile models toward workstation-scale capability in a pocket.

Key facts

Fact Value
Model Bonsai 27B (PrismML), derived from Qwen3.6-27B
Ternary variant ~5.9 GB, 1.71 bpw, ~94.6% of FP16
1-bit variant ~3.9 GB, 1.125 bpw, ~89.5% of FP16
Phone demo ~11 tok/s on iPhone 17 Pro Max (MLX Swift)
Context & capabilities ~262K context; multimodal; reasoning, tools, agentic use
License & date Apache 2.0; announced July 14, 2026

Details

Bonsai 27B is PrismML’s low-bit multimodal flagship, built on Alibaba’s Qwen3.6-27B without architectural changes. The core innovation is extreme quantization: weights are forced to ternary or binary values so a 27B-class model fits under phone and tight laptop memory budgets that conventional packs cannot clear.

That memory cut is why the release matters. Phone apps face hard RAM limits—on iOS often around 6 GB per app—so ordinary 27B builds never shipped on-device. The 1-bit Bonsai variant is promoted as the first of its class on a consumer smartphone, with publisher-reported speeds of about 11 tokens/s on an iPhone 17 Pro Max via MLX Swift. Atomic Chat demonstrated the same open weights live for everyday local chat on iPhone and Android.

Everything ships under Apache 2.0 on Hugging Face in GGUF, MLX, and related formats, with specialized llama.cpp and MLX runtimes for the unusual bit packing. Ternary aims at laptop-class quality; 1-bit aims at phone footprint. Together they move private, offline AI from small mobile models toward something closer to workstation-scale capability in a pocket.

Sources