If AI lab PrismML isn’t in your radar but, it must be — not as a result of it’s raised gobs of cash (it hasn’t but, only a $22.25 million seed spherical), however due to the technical minds concerned and the doubtless industry-changing tech it’s creating.
PrismML is betting that succesful, high-performing, reasoning giant language fashions don’t, in reality, must be giant.
It’s making reasoning fashions so small they will match on PCs and smartphones. (It’s even rumored to be in talks with Apple, although CEO Babak Hassibi declined to touch upon that to TechCrunch.)
On Thursday, PrismML launched Bonsai 2 27B, its newest in a household of fashions, which compresses Qwen3.8 27B, a broadly used open supply mannequin from Alibaba, down to five.9 GB. That’s sufficiently small to suit on a PC and, presumably, a high-end smartphone. It’s a 9x to 10x discount in reminiscence versus the unique.
PrismML was based by a gaggle of Caltech researchers and is led by Hassibi, a Caltech professor and an knowledgeable in compression applied sciences. The startup additionally counts Ion Stoica as an adviser. Stoica is a co-founder of Databricks (and different corporations) and the director of Berkeley’s famed Sky Computing Lab, which has birthed many applied sciences and startups, from Letta to SGLang.
PrismML can also be backed by traders Khosla Ventures, Cerberus Capital, and Caltech.
This startup is definitely not the one firm engaged on LLM compression tech. Multiverse Computing, based by a well known professor from Spain’s Donostia Worldwide Physics Middle, is one other. (And Multiverse Computing has raised gobs of money.)
However Hassibi says that PrismML’s compression tech is exclusive as a result of its LLMs have misplaced nearly no efficiency in contrast with the originals. Bonsai 2 matches 98% of Qwen’s mixture benchmark scores. That’s up from the primary Bonsai, launched a few months in the past in March, that matched 95%. That unique mannequin has already been downloaded over 11 million occasions, and PrismML’s even smaller fashions have been downloaded one other 2.6 million occasions, the corporate says.
So this exhibits that PrismML’s compression outcomes have improved from one launch to the following. Whether or not it might ever get to 100% benchmark efficiency parity is a query that is still to be seen. Compression will possible at all times have some influence, Hassibi says.
Nonetheless, excellent benchmark parity is pretty tutorial anyway. LLMs should not so correct of their uncompressed kind, and benchmarks not so completely reflective of precise duties, {that a} 2% degradation would possible meaningfully have an effect on how a mannequin performs in precise use. (Plus, the encircling software program — the harness a mannequin runs within — issues quite a bit relating to accuracy, too.)
PrismML says it achieves this by shrinking the “weights” that make up a mannequin — weights are, primarily, the knowledge a mannequin learns and shops throughout coaching. Usually, every weight requires 16 bits. PrismML’s strategy, known as “ternary” weights, simplifies that down to 3: +1, −1, or 0. With far smaller values to retailer for every weight, the mannequin takes up dramatically much less area. (For a deeper dive on the compression method, right here’s the venture’s Hugging Face web page.)
The startup’s subsequent aim is to use this compression method to even greater fashions. “The following fashions that we are going to launch, hopefully within the subsequent couple of months, shall be within the several-hundred-billion-parameter vary, and I anticipate it is going to be simpler to retain the intelligence there,” Hassibi instructed TechCrunch.
As mannequin dimension grows, he added, “There may be extra room to have the ability to compress them with out dropping the intelligence. So I’d simply say, as a common development, for bigger fashions, it’s simpler to get to 100%.”
Stoica tells us that he’s excited for this tech as a result of it’s making it doable for superior fashions to run on customers’ units. “You will have intelligence at your fingertips, and it’s going to be free as a result of it’s going to run on the machine you already purchased. It’s additionally going to be personal, since you’re not going to ship it to the cloud.”
While you buy by hyperlinks in our articles, we might earn a small fee. This doesn’t have an effect on our editorial independence.




