跳到正文
北京时间
原文
蚂蚁 inclusionAI:GitHub 新仓库· inclusionAI·· 2026-04-20精选AI 评分69

DR-Venus:基于开放数据的边缘级深度研究智能体

inclusionAI/DR-Venus

AI 导读

DR-Venus 是一个仅用1万条开放数据训练的40亿参数深度研究智能体,基于Qwen3-4B-Thinking-2507架构,支持200步工具调用和超20万tokens的上下文。它通过监督微调与强化学习两阶段训练,在BrowseComp、GAIA等多个深度研究基准上树立了小模型性能新标杆。其SFT版本已超越多数同类开源模型,而RL版本进一步将长程任务可靠性和工具使用校准度提升2-3个百分点。项目已全面开源模型、代码与训练流程。

推荐理由

4B 参数、仅用 1 万条公开数据就能在多个 deep research benchmark 上碾压 8B 对手,蚂蚁 inclusionAI 这次证明了小模型做 Agent 的关键不在参数量而在数据管线,做端侧 Agent 的团队值得拆一下它的 SFT+RL 流程。

正文 · 原文

DR-Venus is a 4B-parameter deep research agent trained entirely on open data. It establishes a new small-model frontier on multiple deep research benchmarks, demonstrating that strong agentic capabilities can emerge from careful data curation and effective training strategies at edge scale.

DR-Venus overview figure

Figure 1: Overview of the DR-Venus training pipeline.

Backbone Qwen3-4B-Thinking-2507
Training Data Open-data only (REDSearcher Data)
Tool Protocol search + visit
Interaction Horizon Up to 200 tool-call steps
Context Length 200K+ (training) / 256K (inference)

🔥 News

📖 Overview

The core goal of DR-Venus is to build a strong edge-scale deep research agent under limited open-data supervision by improving both data quality and effective data utilization. The project consists of three stages:

Stage Description Code
1. SFT Convert raw REDSearcher trajectories into a unified agent format, clean noisy tool interactions, filter for correctness, and upweight long-horizon traces via turn-aware resampling before supervised fine-tuning. SFT/
2. RL Starting from the SFT checkpoint, apply long-horizon reinforcement learning with IGPO-style information gain rewards and turn-level format-aware penalties. RL/
3. Inference Deploy the trained model with the same search + visit tool protocol used during training. Inference/

📊 Main Results

DR-Venus-4B establishes a strong small-model frontier on multiple deep research benchmarks.

Comparison with Small Open Models

Model BrowseComp BrowseComp-ZH GAIA (Text) xBench-DS-2505 xBench-DS-2510 DeepSearchQA
DeepDive-9B-SFT 5.6 15.7 -- 35.0 -- --
DeepDive-9B-RL 6.3 15.1 -- 38.0 -- --
WebSailor-7B 6.7 14.2 37.9 34.3 -- --
OffSeeker-8B-SFT 10.6 24.2 47.6 48.0 -- --
OffSeeker-8B-DPO 12.8 26.6 51.5 49.0 -- --
WebExplorer-8B-RL 15.7 32.0 50.0 53.7 23.0 17.8
AgentCPM-Explore-4B 24.1 29.1 63.9 70.0 34.0 32.8
DR-Venus-4B-SFT 26.8 35.7 65.4 69.0 35.3 37.7
DR-Venus-4B-RL 29.1 37.7 64.4 74.7 40.7 39.6

Key takeaways:

  • DR-Venus-4B-SFT already outperforms prior small open agents on most tracked benchmarks.
  • DR-Venus-4B-RL further improves over SFT by +2.3 on BrowseComp and +2.0 on BrowseComp-ZH.
  • RL mainly improves long-horizon execution reliability, formatting stability, and tool-use calibration.

DR-Venus Pass@K results on BrowseComp and BrowseComp-ZH

Figure 2: Pass@K comparison on BrowseComp and BrowseComp-ZH.

📦 Data Pipeline

The SFT pipeline is built from open REDSearcher trajectories and focuses on making limited supervision more useful for a small model.

Raw REDSearcher Trajectories (10,001)
  │
  ├── Structural Cleaning    ── environment alignment, tool normalization,
  │                              disallowed-tool pruning, duplicate removal
  ├── Correctness Filtering  ── keep trajectories with correct final answers
  │                              → 9,365 trajectories
  └── Turn-Aware Resampling  ── upweight longer trajectories to emphasize
                                 deep research behavior
                                 → 18,745 final SFT instances

For RL, we use open QA supervision curated from the REDSearcher RL data source.

✨ Demo

DR-Venus Demo Video

DeepResearch on problem:"《真探》第一季里有一句类似“useless spin”的台词出现在第几集?".

DR-Venus Demo

🚀 Quick Start

Each subdirectory has its own dependencies and entry scripts. Refer to the subproject READMEs for full details.

1. Inference

Run the trained model with search and visit tools:

cd Inference
pip install -r requirements.txt

# Configure API credentials used by the tool server first
bash run_demo.sh

# Or launch the web demo
bash run_web_demo.sh --model_path <MODEL_PATH> --num_gpus <NUM_GPUS>

See Inference/README.md for the full setup guide.

2. SFT

Prepare cleaned trajectories and train the SFT checkpoint:

cd SFT
pip install -e .
pip install -r requirements.txt

# Optional: convert raw RED trajectories into SFT-ready parquet
python data_clean/prepare_trajectories.py \
  --input <RAW_PARQUET_OR_DIR> \
  --output-dir <OUTPUT_DIR>

# Run supervised fine-tuning
bash train_sft.sh

See SFT/README.md for data format, environment variables, and checkpoint merging.

3. RL

Continue from the SFT checkpoint with long-horizon RL:

cd RL
pip install -r requirements.txt

# Edit .env and train_igpo.sh first
bash train_igpo.sh

See RL/README.md for reward configuration, rollout setup, and troubleshooting.

🔌 External Services

Both Inference/ and RL/ depend on external tools and model endpoints. SFT does not require these.

Service Purpose
Serper Web search
Jina Reader Webpage fetching
OpenAI-compatible API Page summarization
OpenAI-compatible judge model RL reward evaluation

📁 Repository Layout

DR-Venus/
├── README.md
├── assets/
│   ├── dr-venus-overview.png
│   └── dr-venus-passk.png
├── SFT/
│   ├── data_clean/          # RED trajectory conversion & cleaning
│   ├── scripts/             # Checkpoint merge utilities
│   ├── sft_shells/
│   ├── train_sft.sh
│   └── README.md
├── RL/
│   ├── configs/
│   ├── data/
│   ├── tool_server/
│   ├── train_igpo.sh
│   └── README.md
├── Inference/
│   ├── tool_server/
│   ├── run_demo.sh
│   ├── run_web_demo.sh
│   └── README.md
└── model_cards/

🤝 Acknowledgements

DR-Venus builds on several strong open-source foundations:

📝 Citation

@article{venus2026drvenus,
  title={DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data},
  author={Venus Team and Dai, Sunhao and Deng, Yong and Lin, Jinzhen and Song, Yusheng and Wang, Guoqing and Wu, Xiaofeng and Zhou, Yuqi and Yang, Shuo and Ying, Zhenzhe and Zhang, Zhanwei and Meng, Changhua and Wang, Weiqiang},
  journal={arXiv preprint arXiv:2604.19859},
  year={2026}
}

来源:蚂蚁 inclusionAI:GitHub 新仓库 · github.com