跳到正文
北京时间
原文
OpenBMB· @OpenBMB · X·· 2026-06-01精选AI 评分78
AI 导读

OpenBMB联合清华NLP与Modelbest发布两个开源数据集:Ultra-FineWeb-L3(预训练合成数据)包含600B+ tokens(超400B英文、200B+中文),是迄今最大开源中文预训练合成数据集;UltraData-SFT-2605(后训练SFT数据)包含15M+样本,是中国首个开源且包含思考与非思考标注的大规模SFT数据集,覆盖数学、代码、知识和指令遵循。两者均基于UltraData L0-L4框架构建,并在MiniCPM5-1B训练中完成验证。数据集已在HuggingFace免费开放。

推荐理由

面壁开源了两个王炸数据集,预训练的 600B+ token 中文合成数据史上最大,SFT 那边 1500 万条带思考链的指令更是头一回见,做中文基础模型的可以无脑下载了。

正文 · 原文

🏆 Big news! UltraData just hit #1 AND #2 on HuggingFace Trending worldwide! 🎉
Released by OpenBMB × @TsinghuaNLP × Modelbest — two massive open-source datasets now free for everyone:
🔥 Ultra-FineWeb-L3 (web pretraining synthetic data) → 600B+ tokens (400B+ English, 200B+ Chinese) → Largest open-source Chinese pretraining synthetic dataset to date → Built to maximize learnability per token
🔥 UltraData-SFT-2605 (post-training SFT data) → China's first open-source 15M+ SFT dataset with both thinking & non-thinking annotations → Covers math, code, knowledge & instruction-following → Fully traceable data pipeline
🧱 Both built on the UltraData L0–L4 five-tier data management framework, validated end-to-end on MiniCPM5-1B training.
Free to download now 👇
https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3
https://huggingface.co/datasets/openbmb/UltraData-SFT-2605
#OpenSource #LLM #AI #HuggingFace #MiniCPM #UltraData

来源:OpenBMB · x.com