Tech Meridian ← LIVE FEED
RU

RELEASE · MODELS · #582

PrismML releases Bonsai 2 27B — 9–10× compressed Qwen3.8 27B (5.9 GB)

PrismML released Bonsai 2 27B, a compressed version of Alibaba’s open-source Qwen3.8 27B reduced to about 5.9 GB (a 9–10× memory reduction) while matching roughly 98% of Qwen’s aggregate benchmark scores. The Caltech-founded startup says it uses a ternary-weights compression technique to shrink model weights, claims minimal performance loss versus originals, has previously seen millions of downloads of earlier Bonsai releases, and plans to apply the approach to much larger models; it raised a $22.25M seed and is backed by investors including Khosla Ventures and Cerberus Capital.

KEY POINTS

  1. PrismML released Bonsai 2 27B, a compressed version of Alibaba’s open-source Qwen3.8 27B reduced to about 5.9 GB (a 9–10× memory reduction) while matching roughly 98% of Qwen’s aggregate benchmark scores.
  2. The Caltech-founded startup says it uses a ternary-weights compression technique to shrink model weights, claims minimal performance loss versus originals, has previously seen millions of downloads of earlier Bonsai releases, and plans to apply the approach to much larger models; it raised a $22.25M seed and is backed by investors including Khosla Ventures and Cerberus Capital.
  3. Smaller, near-parity LLMs that can run on PCs and phones would change deployment, cost, latency, and privacy trade-offs for many AI applications.

WHY IT MATTERS

Smaller, near-parity LLMs that can run on PCs and phones would change deployment, cost, latency, and privacy trade-offs for many AI applications.

SOURCES & TIMELINE

1