Product Launches

PrismML's 5.9GB Bonsai 2 27B model matches 98% of Qwen's benchmarks

Share
PrismML's 5.9GB Bonsai 2 27B model matches 98% of Qwen's benchmarks

PrismML has released Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B model that matches 98% of its benchmarks while being small enough to run on PCs and high-end smartphones. The startup, backed by notable investors and advisors, aims to make advanced AI models more accessible and private.

TL;DR

  • PrismML's Bonsai 2 27B model reduces Qwen3.8 27B's size to 5.9GB, enabling deployment on local devices.
  • The model matches 98% of Qwen's aggregate benchmark scores, improving from the previous 95% in Bonsai 1.
  • PrismML's compression technique, using ternary weights, significantly reduces model size with minimal performance loss.

What happened

PrismML, an AI lab founded by Caltech researchers, has released Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B model. The new model is just 5.9GB, a 9x to 10x reduction in memory, making it suitable for PCs and high-end smartphones. According to PrismML, Bonsai 2 27B matches 98% of Qwen's aggregate benchmark scores, up from 95% in the previous Bonsai 1 model released in March. The original Bonsai model has been downloaded over 11 million times, with an additional 2.6 million downloads for PrismML's even smaller models. PrismML's compression technique, called ternary weights, simplifies the model's weights to +1, -1, or 0, dramatically reducing the model's size. The startup is backed by investors such as Khosla Ventures, Cerberus Capital, and Caltech, and counts Ion Stoica, a co-founder of Databricks, as an advisor.

Why it matters

PrismML's compression technology enables advanced AI models to run on local devices, making AI more accessible and private. By reducing the model size with minimal performance loss, PrismML aims to democratize AI, allowing users to run powerful models on their personal devices without relying on cloud services. This approach also addresses privacy concerns, as data processing can be done locally. However, the startup acknowledges that perfect benchmark parity may not be achievable, and the surrounding software also plays a significant role in model accuracy. PrismML plans to apply its compression technique to even larger models in the coming months, aiming to retain more intelligence as model size increases.

Key facts

  • PrismML's Bonsai 2 27B model compresses Qwen3.8 27B to 5.9GB, a 9x to 10x reduction in memory.
  • Bonsai 2 27B matches 98% of Qwen's aggregate benchmark scores, up from 95% in Bonsai 1.
  • The original Bonsai model has been downloaded over 11 million times, with an additional 2.6 million downloads for smaller models.
  • PrismML's compression technique uses ternary weights, simplifying model weights to +1, -1, or 0.
  • The startup is backed by Khosla Ventures, Cerberus Capital, and Caltech, and counts Ion Stoica as an advisor.
  • PrismML plans to release compressed models in the several-hundred-billion-parameter range in the next couple of months.
  • CEO Babak Hassibi is a Caltech professor and expert in compression technologies.

Context

PrismML is not the only company working on LLM compression technology. Multiverse Computing, founded by a professor from Spain's Donostia International Physics Center, is another notable player in this space. However, PrismML's unique approach and impressive benchmark scores set it apart. The startup's technology has the potential to revolutionize the way we use AI, making it more accessible, private, and efficient. As AI models continue to grow in size and complexity, compression techniques like those developed by PrismML will play a crucial role in enabling widespread adoption and practical applications.

Topics

Join the discussion

Have a take on this story? Weigh in with our community on Facebook.

💬 Discuss on Facebook →