AI modelsSeptember 17, 2026
-40
🧊

PrismML Squeezes 27B Model Into 5.9 GB

PrismML has compressed a 27-billion-parameter model to 5.9 GB while retaining 98 percent of benchmark performance.

#PrismML#LLM#Compression#OnDevice#Qwen
PrismML quetscht 27B-Modell auf 5,9 GB
Share Article
šŸ”„ What happened PrismML, a Caltech spin-off with just $22.25M in seed funding, dropped Bonsai 2 27B. The model compresses Alibaba's Qwen3.8 27B from roughly 54 GB down to 5.9 GB — a 9–10x memory reduction — while hitting 98% of the original's benchmark scores. šŸ’” Why it matters The trick: ternary weights (+1, āˆ’1, 0) instead of 16 bits per parameter. That means a capable reasoning model runs locally on a PC or high-end phone — no cloud, no API bills, no data leaving the device. The first Bonsai already racked up 11M+ downloads. ⚔ Our take If PrismML closes the last 2% gap, cloud inference for everyday reasoning becomes optional. The big labs should either acquire this team or start sweating.
The title, summary and analysis of this item were produced automatically by an AI system and have not been editorially reviewed. They may contain errors, bias or omissions — when in doubt, read the linked original source.