The release marks a significant milestone in the company’s push to migrate complex AI workflows away from cloud-dependent infrastructure. While the original Bonsai 27B iteration demonstrated the viability of local execution, the new version narrows the quality gap between compressed edge models and their full-scale counterparts. Testing across 20 benchmarks—covering math, coding, and vision—yielded an aggregate score of 83.9, compared to the 85.4 managed by the uncompressed version.
PrismML Shrinks 27B Model for Local Edge Deployment
Pasadena-based PrismML has unveiled Bonsai 2 27B, a multimodal language model that crams advanced reasoning and agentic capabilities into a 5.9 GB package. By utilizing ternary compression, the company claims the model retains over 98% of the performance of the full-precision Qwen3.8 base while slashing memory requirements by ninefold.
Babak Hassibi, Founder and CEO of PrismML, positioned the launch as a solution for demanding tasks like long-horizon execution and multimodal understanding that previously required massive server clusters. Optimized for consumer-grade hardware, the model can achieve speeds of up to 143 tokens per second when running on an NVIDIA GeForce RTX 5090. Developers can access the technology under an Apache 2.0 license, reflecting PrismML’s broader strategy to maximize intelligence density per unit of energy and memory.



Comments (0)
No comments yet. Be the first!