DR-Venus:基于开放数据的边缘级深度研究智能体
inclusionAI/DR-Venus
DR-Venus 是一个仅用1万条开放数据训练的40亿参数深度研究智能体,基于Qwen3-4B-Thinking-2507架构,支持200步工具调用和超20万tokens的上下文。它通过监督微调与强化学习两阶段训练,在BrowseComp、GAIA等多个深度研究基准上树立了小模型性能新标杆。其SFT版本已超越多数同类开源模型,而RL版本进一步将长程任务可靠性和工具使用校准度提升2-3个百分点。项目已全面开源模型、代码与训练流程。
4B 参数、仅用 1 万条公开数据就能在多个 deep research benchmark 上碾压 8B 对手,蚂蚁 inclusionAI 这次证明了小模型做 Agent 的关键不在参数量而在数据管线,做端侧 Agent 的团队值得拆一下它的 SFT+RL 流程。
DR-Venus is a 4B-parameter deep research agent trained entirely on open data. It establishes a new small-model frontier on multiple deep research benchmarks, demonstrating that strong agentic capabilities can emerge from careful data curation and effective training strategies at edge scale.
Figure 1: Overview of the DR-Venus training pipeline.
| Backbone | Qwen3-4B-Thinking-2507 |
| Training Data | Open-data only (REDSearcher Data) |
| Tool Protocol | search + visit |
| Interaction Horizon | Up to 200 tool-call steps |
| Context Length | 200K+ (training) / 256K (inference) |
🔥 News
2026-04-24Released model checkpoints of GGUF version:DR-Venus-4B-SFT-GGUFandDR-Venus-4B-RL-GGUF.2026-04-23Our technical report is now available on arXiv and HF Daily Paper (#3 Paper of the day).2026-04-22Released model checkpoints:DR-Venus-4B-SFTandDR-Venus-4B-RL.2026-04-22Open-sourced the full training and inference codebase.
📖 Overview
The core goal of DR-Venus is to build a strong edge-scale deep research agent under limited open-data supervision by improving both data quality and effective data utilization. The project consists of three stages:
| Stage | Description | Code |
|---|---|---|
| 1. SFT | Convert raw REDSearcher trajectories into a unified agent format, clean noisy tool interactions, filter for correctness, and upweight long-horizon traces via turn-aware resampling before supervised fine-tuning. | SFT/ |
| 2. RL | Starting from the SFT checkpoint, apply long-horizon reinforcement learning with IGPO-style information gain rewards and turn-level format-aware penalties. | RL/ |
| 3. Inference | Deploy the trained model with the same search + visit tool protocol used during training. | Inference/ |
📊 Main Results
DR-Venus-4B establishes a strong small-model frontier on multiple deep research benchmarks.
Comparison with Small Open Models
| Model | BrowseComp | BrowseComp-ZH | GAIA (Text) | xBench-DS-2505 | xBench-DS-2510 | DeepSearchQA |
|---|---|---|---|---|---|---|
| DeepDive-9B-SFT | 5.6 | 15.7 | -- | 35.0 | -- | -- |
| DeepDive-9B-RL | 6.3 | 15.1 | -- | 38.0 | -- | -- |
| WebSailor-7B | 6.7 | 14.2 | 37.9 | 34.3 | -- | -- |
| OffSeeker-8B-SFT | 10.6 | 24.2 | 47.6 | 48.0 | -- | -- |
| OffSeeker-8B-DPO | 12.8 | 26.6 | 51.5 | 49.0 | -- | -- |
| WebExplorer-8B-RL | 15.7 | 32.0 | 50.0 | 53.7 | 23.0 | 17.8 |
| AgentCPM-Explore-4B | 24.1 | 29.1 | 63.9 | 70.0 | 34.0 | 32.8 |
| DR-Venus-4B-SFT | 26.8 | 35.7 | 65.4 | 69.0 | 35.3 | 37.7 |
| DR-Venus-4B-RL | 29.1 | 37.7 | 64.4 | 74.7 | 40.7 | 39.6 |
Key takeaways:
DR-Venus-4B-SFTalready outperforms prior small open agents on most tracked benchmarks.DR-Venus-4B-RLfurther improves over SFT by +2.3 on BrowseComp and +2.0 on BrowseComp-ZH.- RL mainly improves long-horizon execution reliability, formatting stability, and tool-use calibration.
Figure 2: Pass@K comparison on BrowseComp and BrowseComp-ZH.
📦 Data Pipeline
The SFT pipeline is built from open REDSearcher trajectories and focuses on making limited supervision more useful for a small model.
Raw REDSearcher Trajectories (10,001)
│
├── Structural Cleaning ── environment alignment, tool normalization,
│ disallowed-tool pruning, duplicate removal
├── Correctness Filtering ── keep trajectories with correct final answers
│ → 9,365 trajectories
└── Turn-Aware Resampling ── upweight longer trajectories to emphasize
deep research behavior
→ 18,745 final SFT instances
For RL, we use open QA supervision curated from the REDSearcher RL data source.
✨ Demo
DR-Venus Demo Video
DeepResearch on problem:"《真探》第一季里有一句类似“useless spin”的台词出现在第几集?".

🚀 Quick Start
Each subdirectory has its own dependencies and entry scripts. Refer to the subproject READMEs for full details.
1. Inference
Run the trained model with search and visit tools:
cd Inference
pip install -r requirements.txt
# Configure API credentials used by the tool server first
bash run_demo.sh
# Or launch the web demo
bash run_web_demo.sh --model_path <MODEL_PATH> --num_gpus <NUM_GPUS>
See
Inference/README.mdfor the full setup guide.
2. SFT
Prepare cleaned trajectories and train the SFT checkpoint:
cd SFT
pip install -e .
pip install -r requirements.txt
# Optional: convert raw RED trajectories into SFT-ready parquet
python data_clean/prepare_trajectories.py \
--input <RAW_PARQUET_OR_DIR> \
--output-dir <OUTPUT_DIR>
# Run supervised fine-tuning
bash train_sft.sh
See
SFT/README.mdfor data format, environment variables, and checkpoint merging.
3. RL
Continue from the SFT checkpoint with long-horizon RL:
cd RL
pip install -r requirements.txt
# Edit .env and train_igpo.sh first
bash train_igpo.sh
See
RL/README.mdfor reward configuration, rollout setup, and troubleshooting.
🔌 External Services
Both Inference/ and RL/ depend on external tools and model endpoints. SFT does not require these.
| Service | Purpose |
|---|---|
| Serper | Web search |
| Jina Reader | Webpage fetching |
| OpenAI-compatible API | Page summarization |
| OpenAI-compatible judge model | RL reward evaluation |
📁 Repository Layout
DR-Venus/
├── README.md
├── assets/
│ ├── dr-venus-overview.png
│ └── dr-venus-passk.png
├── SFT/
│ ├── data_clean/ # RED trajectory conversion & cleaning
│ ├── scripts/ # Checkpoint merge utilities
│ ├── sft_shells/
│ ├── train_sft.sh
│ └── README.md
├── RL/
│ ├── configs/
│ ├── data/
│ ├── tool_server/
│ ├── train_igpo.sh
│ └── README.md
├── Inference/
│ ├── tool_server/
│ ├── run_demo.sh
│ ├── run_web_demo.sh
│ └── README.md
└── model_cards/
🤝 Acknowledgements
DR-Venus builds on several strong open-source foundations:
- verl -- training infrastructure
- IGPO -- RL algorithmic foundation
- Tongyi DeepResearch -- deep research agent design references
- REDSearcher -- open-data trajectories and related tooling
📝 Citation
@article{venus2026drvenus,
title={DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data},
author={Venus Team and Dai, Sunhao and Deng, Yong and Lin, Jinzhen and Song, Yusheng and Wang, Guoqing and Wu, Xiaofeng and Zhou, Yuqi and Yang, Shuo and Ying, Zhenzhe and Zhang, Zhanwei and Meng, Changhua and Wang, Weiqiang},
journal={arXiv preprint arXiv:2604.19859},
year={2026}
}
来源:蚂蚁 inclusionAI:GitHub 新仓库 · github.com