Skip to main content

PrismML unveils Bonsai 2 27B model designed for PCs

Caltech-backed PrismML launches Bonsai 2 27B, compressing Alibaba's Qwen3.8 27B down to 5.9 GB to run on PCs and smartphones.

AI-written
Inewgen
18 Sep 2026Source: TechCrunch3 min read (0 views)
Share
PrismML unveils Bonsai 2 27B model designed for PCs

Stock photo for illustration only, not from the actual event

Font size
  • PrismML has released Bonsai 2 27B, a compressed reasoning model sized at just 5.9 GB.
  • The model is compact enough to run locally on personal computers and high-end smartphones.
  • It achieves a 9x to 10x memory reduction compared to Alibaba's original Qwen3.8 27B model.
  • The startup was founded by Caltech researchers and is led by compression technology expert Babak Hassibi.

As the artificial intelligence industry continues to rely heavily on massive data centers for running complex models, a new startup called PrismML is betting that high-performing reasoning models do not actually need to be large. The company is actively shrinking these models so they can fit directly onto personal computers and smartphones.

On Thursday, PrismML released Bonsai 2 27B, the latest addition to its model family. The model compresses Qwen3.8 27B, a widely used open-source model from Alibaba, down to just 5.9 gigabytes. This footprint is small enough to run on a standard PC and potentially a high-end smartphone, representing a 9x to 10x reduction in memory requirements compared to the original architecture.

5.9 GBFile size of the Bonsai 2 27B model
9-10xMemory reduction versus original model
27BParameter count baseline

PrismML was founded by a team of researchers from the California Institute of Technology and is led by Babak Hassibi, a Caltech professor and expert in compression technologies. The startup also counts Ion Stoica, co-founder of Databricks and director of Berkeley's Sky Computing Lab, among its advisors. Financial backing has been provided by prominent investors including Khosla Ventures, Cerberus Capital, and Caltech itself.

modern minimalist contemporary architecture

Stock photo for illustration only, not from the actual event

Large language model compression has emerged as a critical field in AI development. Running massive models entirely on cloud infrastructure incurs substantial ongoing costs and introduces latency. Deploying compressed models directly to edge devices improves user privacy and enables offline functionality, though engineers must carefully balance file size reductions against potential degradations in reasoning accuracy.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

PrismML is not alone in pursuing LLM compression technology. Multiverse Computing, founded by a prominent professor from Spain's Donostia International Physics Center, is another player in the space and has successfully secured significant funding, highlighting strong investor interest in efficient AI deployment.

"Compression will likely always have some impact."

Babak Hassibi, CEO of PrismML

While achieving 100 percent benchmark parity remains an ongoing question, PrismML's leadership notes that minor performance trade-offs are often negligible in practical application. Uncompressed models are imperfect to begin with, and standard benchmarks rarely reflect real-world tasks completely. Furthermore, the surrounding software harness plays a crucial role in maintaining overall accuracy during everyday use.

Source: TechCrunch

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article