ModelBest、清华大学与OpenBMB社区联合发布了BitCPM-CANN,这是全球首个完全基于华为昇腾910B NPU训练的开源1.58比特三元大模型。其核心创新在于采用仅含三种权重状态的极低比特量化技术,使模型内存占用相比BF16降低约6倍,可高效部署于手机、电脑、车载设备等边缘端。更关键的是,整个训练全栈(从量化算子到框架)均在昇腾上原生构建与验证,而非简单移植。该模型家族(0.5B-8B)在多项基准测试上保持了全精度模型95-97%的性能,为资源受限环境下部署和复现大模型提供了可落地的解决方案。
首个开源的1.58-bit三元LLM,直接在昇腾芯片上原生训练,内存压缩到BF16的六分之一,8B模型就能跑在手机上,做端侧部署的可以立刻上手试试了。
BitCPM-CANN just became the world’s first open-sourced 1.58-bit ternary LLM trained entirely on Chinese-developed AI infrastructure.
Developed by ModelBest, Tsinghua Univ, and OpenBMB community, the entire training pipeline, from quantization operators and algorithms to the full-stack framework, was natively executed on Huawei Ascend 910B NPUs.
1.58-bit ternary weights use only 3 weight states, so the model needs far less memory when deployed on phones, PCs, cars, and local industrial devices.
The harder achievement is the training system behind it: QAT, STE, low-bit operators, algorithms, framework work, and reproducible training scripts all had to hold together on Ascend 910B.
When hardware costs rise, the winning model is not merely the one that scores higher in a chart, but the one that can be trained, reproduced, deployed, and improved under real constraints.
🚀 BitCPM-CANN by ModelBest × @Tsinghua_Uni × OpenBMB is here — and it's not about stacking parameters. Memory costs are skyrocketing. Hardware constraints are tightening. Edge AI needs smarter solutions — and BitCPM-CANN delivers!🎉 ✅ Edge-ready: 8B model runs smoothly on mobile, PC, and automotive devices. Combined with MoE, even 100B->60B scale models could fit on terminal hardware. ✅ Memory-efficient: ~6× lower memory footprint vs. BF16 — unlock significantly more model capacity without adding physical RAM. No new chips required. ✅ Natively built on Ascend: The first 1.58-bit training pipeline completed end-to-end on Huawei Ascend 910B — from quantization kernels to the full training stack. Not a port. Built natively on Ascend from day one. ✅ Full model family, fully verified, 0.5B–8B: Each model is fully aligned with its full-precision counterpart, covering 11 benchmark tasks with 95–97% (1B-8B) retention compared to full-precision MiniCPM4. Open-source and fully reproducible — from research to deployment, you can run any size confidently. BitCPM-CANN isn’t just a model — it’s the result of years of engineering rigor, turning complex research into something you can actually deploy. Open-source, 0.5B to 8B. Try it now!👇 🤗 Hugging Face: https://huggingface.openbmb.com/collections/openbmb/bitcpm4-cann 🔭 ModelScope: https://www.modelscope.cn/collections/OpenBMB/BitCPM4-CANN在 X 查看被引用的帖子
来源:Rohan Paul · x.com