đ Zaprep: Deine Socials auf Steroiden. Kostenlos starten â automatisiere 1.000 DMs/Monat und verwandle Engagement in Leads. mit 1.000 automatisierten DMs/Monat.
AI NewsPrismML hopes its tiny LLM will change how we all use AI
PrismML hopes its tiny LLM will change how we all use AI
6:26 AM IST ¡ September 18, 2026

If AI labPrismMLisnât on your radar yet, it should be â not because itâs raised gobs of money (it hasnât yet, just a $22.25 million seed round), but because of the technical minds involved and the potentially industry-changing tech itâs developing. PrismML is betting that capable, high-performing, reasoning large language models donât, in fact, have to be large. It is making reasoning models so small they can fit on PCs and smartphones. (Itâs even rumoredto be in talks with Apple, though CEO Babak Hassibi declined to comment on that to TechCrunch.) On Thursday, PrismMLreleased Bonsai 2 27B, its latest in a family of models, which compresses Qwen3.8 27B, a widely used open source model from Alibaba, down to 5.9 GB. Thatâs small enough to fit on a PC and, possibly, a high-end smartphone. Itâs a 9x to 10x reduction in memory versus the original. PrismML was founded by a group of Caltech researchers and is led by Hassibi, a Caltech professor and an expert in compression technologies. The startup also counts Ion Stoica as an adviser. Stoica is a co-founder of Databricks (and other companies) and the director of Berkeleyâs famed Sky Computing Lab, which has birthed many technologies and startups, fromLettatoSGLang. PrismML is also backed by investors Khosla Ventures, Cerberus Capital, and Caltech. This startup is certainly not the only company working on LLM compression tech. Multiverse Computing, founded by a well-known professor from Spainâs Donostia International Physics Center, is another. (And Multiverse Computinghas raised gobs of cash.) But Hassibi says that PrismMLâs compression tech is unique because its LLMs have lost virtually no performance compared with the originals. Bonsai 2 matches 98% of Qwenâs aggregate benchmark scores. Thatâs up from the first Bonsai, released a couple of months ago in March, that matched 95%. That original model has already been downloaded over 11 million times, and PrismMLâs even smaller models have been downloaded another 2.6 million times, the company says. So this shows that PrismMLâs compression results have improved from one release to the next. Whether it could ever get to 100% benchmark performance parity is a question that remains to be seen. Compression will likely always havesomeimpact, Hassibi says. Still, perfect benchmark parity is fairly academic anyway. LLMs are not so accurate in their uncompressed form, and benchmarks not so perfectly reflective of actual tasks, that a 2% degradation would likely meaningfully affect how a model performs in actual use. (Plus, the surrounding software â the harness a model runs inside of âmatters a lot when it comes to accuracy, too.) PrismML says it achieves this by shrinking the âweightsâ that make up a model â weights are, essentially, the information a model learns and stores during training. Normally, each weight requires 16 bits. PrismMLâs approach, called âternaryâ weights, simplifies that down to three: +1, â1, or 0. With far smaller values to store for each weight, the model takes up dramatically less space. (For a deeper dive on the compression technique, hereâs the projectâsHugging Face page.) The startupâs next goal is to apply this compression technique to even bigger models. âThe next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there,â Hassibi told TechCrunch. As model size grows, he added, âThere is more room to be able to compress them without losing the intelligence. So I would just say, as a general trend, for larger models, itâs easier to get to 100%.â Stoica tells us that heâs excited for this tech because itâs making it possible for advanced models to run on usersâ devices. âYou are going to have intelligence at your fingertips, and itâs going to be free because itâs going to run on the device you already bought. Itâs also going to be private, because youâre not going to send it to the cloud.â
read moreAktuelle KI-News
Alle News anzeigen âTool einreichen
PoweredByAI.app ist ein KI-Tools-Verzeichnis, das Menschen, Unternehmen und Creators hilft, die besten KI-Tools fßr Schreiben, Programmieren, Design, Produktivität und mehr zu entdecken.
Š 2026 , Ein Produkt von011BQ. Alle Rechte vorbehalten.




