迈向自我进化的智能文献检索系统
Towards Self-Evolving Agentic Literature Retrieval
针对传统检索无法理解复杂意图、而前沿大语言模型成本高且存在幻觉的问题,研究团队提出了自我进化的智能文献检索系统PaSaMaster。该系统通过迭代式意图分析、检索与排序,将文献检索转变为动态演进的过程,并采用三项关键设计:利用排序证据揭示信息缺口以优化搜索;将检索定义为意图-论文相关性排序任务,从根本上杜绝虚假文献;通过分离规划与检索来提升效率,仅用大模型理解意图,而将大规模检索与评分交由轻量模型处理。在涵盖38个学科的基准测试中,该系统将传统关键词检索的F1分数提升15.6倍,完全消除了文献幻觉,且性能超越GPT-5.2达30%,计算成本仅为后者的1%。
学术文献检索一直被关键词和LLM幻觉两头堵,这个系统用规划与检索分离做到了零幻觉,F1暴涨15.6倍,比GPT-5.2强30%却只花1%算力,做科研的可以马上跑起来。
sjtu0267098@sjtu.edu.cn
\fnm
Tian
\sur
Jin
\fnm
Jing
\sur
Kang
\fnm
Xianghe
\sur
Pang
\fnm
Jingyi
\sur
Chai
\fnm
Tingjia
\sur
Miao
\fnm
Fenyi
\sur
Liu
\fnm
WenHao
\sur
Wang
\fnm
Sikai
\sur
Yao
\fnm
Yuzhi
\sur
Zhang
sihengc@sjtu.edu.cn
Abstract
Scientific literature retrieval must understand complex search intents while preserving source authenticity. Traditional keyword and embedding-based systems return authentic sources but miss nuanced intents, whereas large language models capture richer intents but may fabricate citations. We introduce PaSaMaster, a Recursive Self-Evolving agentic literature retrieval system that iteratively analyzes intent, retrieves verified papers and ranks them with evidence-grounded relevance scores. PaSaMaster combines self-evolving retrieval that refines search intent from ranked evidence over time, hallucination-free ranking over verified papers rather than generated citations, and cost-efficient planning–retrieval separation that reserves frontier LLMs for intent understanding while delegating retrieval and scoring to lightweight models and customized corpora. Across 38 disciplines in PaSaMaster-Bench, PaSaMaster achieves a 16.5 higher F1-score than Google Scholar and a 37.8% higher F1-score than GPT-5.2 at about 1% of the cost, while reducing source hallucination from 32.66% in generative LLMs to zero: https://github.com/sjtu-sai-agents/PaSaMaster
keywords:
Scientific Literature Discovery, Agentic AI, Large Language Models
| Paradigm Level | Representative Systems | Intent Adaptivity | Source Reliability | Cost Efficiency |
| Level 0: Lexical Retrieval | Google Scholar [google_scholar], PubMed [pubmed-ncbi-official] | Keyword-based; severe intent compression | Verified indexed papers | Efficient but semantically shallow |
| Level 1: Semantic Retrieval | OpenScholar [asai2024openscholar], Bohrium Navigator [zhang2025bohriumscimaster] | Passive embedding matching; limited intent compression | Verified indexed papers | Efficient but semantically shallow |
| Level 2: Generative LLMs | GPT-5.2 [openai2026gpt54], Gemini 3.1 Pro [google2026gemini31pro], DeepSeek [deepseekai2025deepseekv32] | Strong natural-language understanding | Prone to hallucinated papers | Expensive to deploy at scale |
| Level 3: Fixed-Pipeline Agentic Retrieval | Google Scholar Labs [google2025scholarlabs], PaSa [he2025pasa] | User intent fixed at the outset; no cognition update during retrieval | Verified indexed papers | Cost-controlled, but constrained by fixed intent interpretation |
| Level 4: Recursive Self-Evolving Agentic Retrieval | PaSaMaster (Ours) | Iteratively refines intent using ranked evidence | Verified indexed papers | Cost-efficient planning–retrieval separation |
a
b
Scientific literature retrieval is the axiomatic starting point of all scientific inquiry [gusenbauer2021searching]. Before formulating hypotheses, designing experiments, or building new theories, researchers must fundamentally navigate the vast and ever-expanding corpus of existing knowledge [fortunato2018science]. However, the volume of scientific publications has grown exponentially over recent decades, decisively overwhelming the fixed cognitive bandwidth of individual researchers [bornmann2015growth, Landhuis_2016, houssard2025gerontocratization]. This severe information overload has driven an inevitable reliance on artificial intelligence to automate and accelerate knowledge discovery [wang2023scientific]. More importantly, modern literature search is rarely a simple keyword lookup [furnas1987vocabulary, beel2010aseo]. Researchers often express complex academic intents involving technical constraints, application contexts, and implicit background knowledge [white2009exploratory, gusenbauer2021searching, ajith2024litsearch]. As large language models reshape scientific research workflows, literature retrieval therefore faces a new central challenge: how to deeply understand complex search intents while ensuring that every returned source is real and verifiable [zhang2025llmscimethod, ajith2024litsearch].
Existing literature retrieval systems still struggle to jointly achieve complex intent understanding and source authenticity. Some methods preserve source authenticity at the cost of shallow intent understanding [google_scholar, asai2024openscholar, zhang2025bohriumscimaster], whereas others improve semantic comprehension while sacrificing factual reliability [deepseekai2025deepseekv32, kimiteam2026kimik2, minimax2025minimaxm1, glm5team2025glm45, google2026gemini31pro, openai2026gpt54]. This creates a persistent tradeoff between reliable but limited retrieval and more intelligent but less trustworthy literature discovery.
The tradeoff between intent understanding and source reliability is clearer when viewed through the evolution of literature retrieval paradigms (Table 1). Level 0 (Lexical Retrieval) [google_scholar, pubmed-ncbi-official] guarantees source authenticity through indexed databases, but reduces complex research intents to rigid keywords, causing severe intent compression. Level 1 (Semantic Retrieval) [asai2024openscholar, zhang2025bohriumscimaster] improves over exact keyword matching by using embedding-based similarity and retrieval-augmented matching [karpukhin2020dense, lewis2021rag], but still treats retrieval as passive query–document matching and lacks the ability to actively clarify, decompose, or refine complex intents. Level 2 (Generative LLMs) [deepseekai2025deepseekv32, kimiteam2026kimik2, minimax2025minimaxm1, glm5team2025glm45, google2026gemini31pro, openai2026gpt54] offers stronger intent comprehension, yet its probabilistic generation introduces fabricated papers [zhang2025sirenssong, Farquhar_2024], undermining the factual trust required for scientific inquiry. Level 3 (Fixed-Pipeline Agentic Retrieval) [google2025scholarlabs, he2025pasa] mitigates source hallucination by grounding LLM agents in verifiable retrieval tools [nakano2021webgpt, schick2023toolformer, qin2023toolllm, yao2023react, du2026openseeker, du2026openseekerv2]. However, these systems typically follow a predefined retrieve–read–answer pipeline in which the interpretation of the user question is largely fixed at the beginning. For complex research intents, this initial interpretation can be incomplete or biased toward only part of the request, causing later retrieval steps to miss relevant subtopics or return papers that satisfy only some constraints. This limitation motivates retrieval systems that can update their understanding of the intent as evidence accumulates.
The unresolved need for adaptive, verifiable and efficient retrieval motivates Level 4 (Recursive Self-Evolving Agentic Retrieval), represented by PaSaMaster. PaSaMaster is a Recursive Self-Evolving agentic literature retrieval system that iteratively analyzes intent, retrieves verified papers and ranks them with evidence-grounded relevance scores. Rather than treating literature search as one-shot query–document matching problem, PaSaMaster formulates scientific literature discovery as a Recursive Self-Evolving intent–paper relevance ranking process. This design enables the system to align with complex research intents while ensuring that every returned source is real, verifiable, and grounded in customized corpora.
PaSaMaster is built on three key designs. First, Recursive Self-Evolving retrieval: it transforms literature retrieval from one-shot query–document matching into an adaptive search process that evolves over time [guo2024multiagent, shinn2023reflexion], where retrieved and ranked evidence is used to identify coverage gaps, refine the research intent, and guide subsequent retrieval rounds. Second, hallucination-free ranking: it treats literature discovery as intent–paper relevance ranking rather than generation, and trains a lightweight Ranker on multidisciplinary query–paper evidence so that Librarian agents can make expert-style relevance judgments across disciplines while ranking only verified papers grounded in original evidence. Third, cost-efficient planning–retrieval separation: it uses frontier LLMs only for intent understanding and refinement, while delegating large-scale retrieval and relevance scoring to customized scientific corpora and lightweight models. Together, these designs enable PaSaMaster to align with complex research intents while maintaining source verifiability and cost-efficient scalability.
To evaluate retrieval capability on complex natural-language literature search problems, we introduce PaSaMaster-Bench, the first multidisciplinary literature retrieval benchmark designed for complex search intents. Unlike conventional retrieval benchmarks [he2025pasa, kang2025researcharenabenchmarkinglargelanguage, bragg2026astabenchrigorousbenchmarkingai, xiong2026autoresearchbenchbenchmarkingaiagents, ajith2024litsearch] built around short keyword queries, PaSaMaster-Bench focuses on highly specific, multi-constrained natural language search intents that require systems to search, verify, and rank all papers satisfying explicit criteria. The benchmark contains 244 expert-curated tasks spanning 38 scientific disciplines, with queries, constraints, target paper lists, and evaluation checklists annotated and verified by human domain experts.
The PaSaMaster-Bench evaluation reveals the severe inaccuracy and incompleteness of traditional keyword retrieval, with PaSaMaster improving F1-score by 16.5. The evaluation also exposes the unreliability of generative LLMs, which exhibit hallucination rates up to 32.66%. Remarkably, PaSaMaster outperforms GPT-5.2 by 37.8% while using only 1% of its computational cost, and maintains zero source hallucination. These results demonstrate that Recursive Self-Evolving, evidence-grounded relevance ranking can improve complex intent understanding, eliminate source hallucination and enable low-cost scientific literature discovery at scale.
1 Results
1.1 PaSaMaster system overview
Figure 1a summarizes how PaSaMaster converts a natural-language search intent into an auditable ranked paper list. The system first makes the implicit structure of the request explicit, turning topical scope, methodological requirements, application context, metadata restrictions and exclusion rules into a retrieval strategy and verification checklist. It then returns top- candidate papers with paper-level relevance scores, recommendation rationales and constraint-level judgments. Each judgment is linked to the evidence used to make it, including metadata, abstracts, full-text snippets and the reasoning connecting the evidence to the stated constraint.
PaSaMaster is organized into three layers. The Navigator performs intent disambiguation, checklist construction, search planning and round-by-round reflection. The Librarian Swarm executes multi-channel retrieval, evidence verification, intent–paper scoring and criteria-based reranking. The verified corpus and toolset layer supplies scientific corpora, search and visit tools, retrieval operators and reading tools for metadata, abstracts and evidence chunks. The strategy–feedback loop between the Navigator and Librarians makes retrieval self-evolving: ranked evidence from each round updates the system’s understanding of the user’s intent and determines the next search direction, while the verified corpus layer keeps all ranked papers authentic and evidence-grounded.
1.2 PaSaMaster-Bench captures complex search intents
PaSaMaster-Bench evaluates whether retrieval systems can resolve the kinds of compositional search intents that arise in real research work [white2009exploratory, gusenbauer2021searching, ajith2024litsearch]. Such intents are rarely reducible to a topic keyword: a user may simultaneously specify the scientific scope, required method, application domain, venue or time window, and exclusions that remove papers that look related but are not the intended target. Because these constraints are expressed in natural language rather than as explicit database filters, the system must infer the full intent before it can identify the correct paper set.
A representative task is: Could you help me find empirical studies on federated learning for multi-hospital collaborative training of medical imaging AI models, published in Nature Medicine, Nature Communications or IEEE Transactions on Medical Imaging between 2023 and October 2025? This is the kind of request a researcher might make when preparing a review, grant proposal or related-work section. It requires the system to satisfy the topic, method, application, venue, time and exclusion constraints jointly; a paper that only discusses federated learning, or only studies medical imaging, is not sufficient.
The benchmark contains 244 independent literature discovery tasks across 38 scientific disciplines. For each task, domain experts write the natural-language intent, decompose it into objective checklist items, assemble a broad candidate pool through multiple retrieval channels and annotate every candidate against the checklist (Fig. 1b). The resulting target set is therefore stricter than topical relevance: a paper is correct only if it satisfies all required constraints. PaSaMaster-Bench measures whether a system can recover the intended paper set behind a realistic research question, not merely whether it can retrieve papers from the same field.
1.3 Experimental setup
We compare PaSaMaster with four groups of baselines. Lexical retrieval systems, represented by Google Scholar [google_scholar], search verified indexed records but mainly rely on keyword matching. Semantic retrieval systems, including OpenScholar [asai2024openscholar] and Bohrium Science Navigator [zhang2025bohriumscimaster], use semantic matching or retrieval-augmented scientific search. Generative LLM baselines include DeepSeek, Kimi, MiniMax, GLM, Gemini and GPT [deepseekai2025deepseekv32, kimiteam2026kimik2, minimax2025minimaxm1, glm5team2025glm45, google2026gemini31pro, openai2026gpt54]. These LLMs are evaluated as tool-assisted search agents rather than closed-book models: each is equipped with Search and Visit tools and prompted to search, inspect evidence and verify sources before returning papers, following tool-use agent evaluation protocols [yao2023react, du2026openseeker, du2026openseekerv2]. For fixed-pipeline agentic retrieval, we use Google Scholar Labs [google2025scholarlabs] as a representative predefined agentic search workflow.
Retrieval quality is measured at using standard information retrieval metrics [manning2008introduction, jarvelin2002cumulated]. Recall@20 measures how many ground-truth target papers are recovered in the top 20 results; Precision@20 measures how many returned top-20 papers are genuine target papers; F1-score@20 summarizes the balance between recall and precision; and NDCG@20 measures whether relevant papers are ranked closer to the top. We also report per-query cost and source hallucination rate. Cost captures the computational expense of interpreting complex constraints and carrying out retrieval, while hallucination rate measures the proportion of returned papers that cannot be matched to a real scholarly record or contain provably wrong source metadata. Together, these metrics assess intent comprehension, source authenticity and cost-effective retrieval.
1.4 Accurate, hallucination-free retrieval at low cost
| Method | NDCG | Recall | Precision | F1-score | Hallucination | Cost ($) |
| Lexical Retrieval Systems | ||||||
| Google Scholar | 2.07 | 1.69 | 1.48 | 1.39 | 0 | – |
| Semantic Retrieval Systems | ||||||
| OpenScholar | 14.61 | 11.68 | 8.52 | 7.92 | 0 | – |
| Bohrium Science Navigator | 22.39 | 19.37 | 12.50 | 12.26 | 0 | – |
| Generative LLMs | ||||||
| DeepSeek-v3.2 | 35.82 | 24.76 | 15.35 | 15.56 | 12.94 | 0.28 |
| Kimi-K2.5 | 37.80 | 28.08 | 16.95 | 17.36 | 26.59 | 0.16 |
| MiniMax-M2.7 | 30.70 | 24.23 | 14.42 | 15.11 | 32.66 | 0.18 |
| GLM-5 | 35.89 | 28.99 | 16.93 | 18.18 | 21.64 | 0.56 |
| Gemini-3.1-pro | 31.34 | 21.30 | 11.68 | 12.48 | 27.54 | 0.38 |
| GPT-5.2 | 31.59 | 25.32 | 16.82 | 16.69 | 5.65 | 6.06 |
| Fixed-Pipeline Agentic Retrieval | ||||||
| Google Scholar Labs | 30.54 | 29.01 | 18.79 | 18.87 | 0 | – |
| Recursive Self-Evolving Agentic Retrieval | ||||||
| PaSaMaster | 39.52 | 33.24 | 23.46 | 23.00 | 0 | 0.05 |
Table 2 reports the main results on PaSaMaster-Bench, and Fig. 2 summarizes cross-disciplinary robustness, source-error patterns and Ranker gains. Overall, PaSaMaster achieves the best retrieval quality while maintaining zero source hallucination and low computational cost. These results support the three central claims of our system: Recursive Self-Evolving retrieval improves understanding of complex natural-language search intents, intent–paper relevance ranking prevents hallucinated sources, and planning–retrieval separation enables cost-efficient large scale literature discovery.
Recursive Self-Evolving retrieval improves complex intent understanding. As shown in Table 2, PaSaMaster achieves the highest retrieval performance across all main quality metrics, with an NDCG@20 of 39.52, Recall@20 of 33.24, Precision@20 of 23.46, and F1-score@20 of 23.00. This demonstrates that PaSaMaster is better able to recover the target papers implied by complex multi-constraint research intents. Compared with Google Scholar [google_scholar], PaSaMaster improves F1-score@20 from 1.39 to 23.00, a 16.5 improvement, showing the severe limitation of keyword-centric retrieval under complex natural-language queries. Compared with semantic retrieval systems, PaSaMaster also substantially outperforms OpenScholar [asai2024openscholar] and Bohrium Science Navigator [zhang2025bohriumscimaster], indicating that passive semantic matching is still insufficient for queries requiring constraint reasoning and intent refinement. Among generative LLMs, the strongest F1-score baseline is GLM-5 [glm5team2025glm45] with 18.18, while the fixed-pipeline agentic retrieval baseline Google Scholar Labs [google2025scholarlabs] achieves 18.87. PaSaMaster reaches 23.00, improving over these strongest baselines by 26.5% and 21.9%, respectively. Fig. 3 provides a mechanistic view of this advantage. In Fig. 3a, papers retrieved in successive rounds shift toward different topic regions, indicating that evidence from earlier rounds changes the system’s interpretation of the query and opens new search directions. In Fig. 3b, the number of recovered ground-truth papers increases across rounds, showing that these new search directions lead to additional relevant papers rather than merely repeating the initial retrieval. Together, the two panels show that Recursive Self-Evolving retrieval improves complex intent understanding by using retrieved evidence to update cognition and guide later searches toward complementary parts of the intended paper set.
Intent–paper relevance ranking eliminates source hallucination. As shown in Table 2, PaSaMaster achieves 0% hallucination while maintaining the strongest retrieval quality. This result directly supports our design choice of treating literature discovery as intent–paper relevance ranking rather than generation. In contrast, the tool-assisted generative LLM baselines [deepseekai2025deepseekv32, kimiteam2026kimik2, minimax2025minimaxm1, glm5team2025glm45, google2026gemini31pro, openai2026gpt54] exhibit substantial hallucination rates, including 32.66% for MiniMax-M2.7, 27.54% for Gemini-3.1, 26.59% for Kimi-K2.5, 21.64% for GLM-5, 12.94% for DeepSeek-v3.2, and 5.65% for GPT-5.2. These results show that even frontier LLMs with search and visit tools [nakano2021webgpt, schick2023toolformer, qin2023toolllm, yao2023react, du2026openseeker, du2026openseekerv2] remain vulnerable to fabricating or misreporting scientific sources [zhang2025sirenssong, Farquhar_2024]. Fig. 2b further shows that hallucinations arise from multiple citation fields, including title, author, date, and link errors. By contrast, PaSaMaster ranks only papers retrieved from verified corpora and grounds relevance judgments in original paper evidence, thereby ensuring zero hallucination in source information.
Planning–retrieval separation reduces cost. Table 2 shows that PaSaMaster achieves this performance at a cost of only $0.05 per query. This is far below GPT-5.2 [openai2026gpt54] at $6.06, GLM-5 [glm5team2025glm45] at $0.56, Gemini-3.1-pro [google2026gemini31pro] at $0.38, and DeepSeek-v3.2 [deepseekai2025deepseekv32] at $0.28. In particular, PaSaMaster outperforms GPT-5.2 in F1-score@20 by 37.8% while using only about 1% of its computational cost. This confirms the benefit of separating high-level planning from large-scale retrieval, enabling PaSaMaster to maintain high-quality retrieval at substantially lower cost.
a
b
c
a
b
Consistent gains across disciplines. Fig. 2 reports retrieval robustness, source hallucination and ranking performance across the 38 scientific disciplines in PaSaMaster-Bench. In Fig. 2a, PaSaMaster achieves leading F1-score distributions across multiple subject groups. Its performance range is shifted upward relative to competing paradigms, with a higher retrieval floor on difficult cases and a higher upper range on easier cases, indicating that the system improves both robustness and peak retrieval quality across disciplines. In Fig. 2c, the overall row shows consistent gains after Ranker training in NDCG@20, Recall@20 and Precision@20, while the subject rows show improvements across most disciplines. These gains reflect training on tens of thousands of multidisciplinary query–paper examples, which helps the Ranker generalize beyond a single domain and place relevant papers more accurately in the final list. Fig. 3b further shows that Recursive Self-Evolving retrieval expands the evidence space round by round across disciplines, increasing the average number of retrieved papers, high-score papers, and recovered ground-truth papers. Together, these results suggest that PaSaMaster’s gains are not driven by a narrow domain-specific advantage. Instead, they reflect the generality of the three design principles: Recursive Self-Evolving retrieval supports complex intent understanding, evidence-grounded ranking ensures source authenticity, and planning–retrieval separation enables scalable retrieval across heterogeneous scientific domains.
2 Discussion
PaSaMaster shows that scientific literature discovery can be framed as a Recursive Self-Evolving, evidence-grounded ranking problem rather than as either keyword matching or citation generation. Across PaSaMaster-Bench, this formulation improves the recovery of target papers under complex search intents while maintaining zero source hallucination and substantially reducing computational cost. This is important because current systems face a structural trade-off: database-backed retrieval preserves source authenticity but often misses the intent behind a nuanced research need, whereas frontier LLMs can reason over richer requests but may fabricate or misreport scientific sources. PaSaMaster narrows this gap by keeping the output space restricted to verified papers while allowing the search process itself to evolve.
The central mechanism is to use ranked evidence to continually update the system’s understanding of the user’s intent during retrieval. This makes PaSaMaster Recursive Self-Evolving: the search process does not merely execute an initial query, but progressively revises what the query means as evidence accumulates. Researchers rarely know the exact best query before seeing the literature; they search, inspect partial results, recognize missing constraints and refine their direction. PaSaMaster operationalizes this process by letting the Navigator revise the retrieval strategy and verification checklist after observing scored candidates from the Librarian swarm. The resulting loop helps the system move beyond a static interpretation of the initial query, which is especially valuable for multi-constraint intents in which relevance depends on combinations of topic, method, dataset, application context and exclusion criteria. The cross-disciplinary results suggest that this mechanism is not only a domain-specific optimization, but a general strategy for aligning retrieval with complex scientific needs.
Equally important is the decision to rank papers rather than generate citations. In many LLM-based literature workflows, hallucination arises because the model is asked to produce bibliographic objects directly from parametric memory or from partially grounded context. PaSaMaster changes this failure mode by requiring every candidate to be retrieved from a verified corpus and every relevance judgment to be supported by evidence from the original paper. The system therefore does not promise that every relevance judgment is perfect, but it does make each relevance score traceable to paper-level evidence and ensures that recommended sources are real and auditable. This distinction matters for scientific use because fabricated papers can distort a literature review, create false evidence chains and mislead subsequent hypothesis formation or experiment design. PaSaMaster reduces this risk by making each recommendation traceable to a real paper and to the evidence used to judge its relevance.
The low-cost design also shapes the potential use of PaSaMaster. Literature discovery is not a one-off task; researchers repeatedly search while designing projects, writing related work, checking novelty and updating reviews. A system that relies on frontier LLMs for every retrieval and scoring operation is difficult to deploy at this frequency and scale. By separating high-level planning from large-scale retrieval and lightweight relevance scoring, PaSaMaster makes agentic literature discovery more practical for broad, multidisciplinary use. In this role, the system should be viewed as an assistive layer that expands and organizes the candidate evidence space, not as a replacement for expert judgment. Its value is strongest when it delivers more comprehensive and accurate retrieval at low cost, exposes why papers were recommended and makes omissions easier to diagnose.
Several limitations remain. First, PaSaMaster-Bench is expert-curated, and its target sets may still be influenced by the disciplinary expertise, prior knowledge and search horizon of the annotators. We mitigated this limitation by using multiple retrieval channels to help experts assemble broad candidate pools, and by requiring elementwise checklist scoring so that each candidate paper is verified against the stated intent rather than accepted by impression alone. Second, the current evaluation focuses on top-ranked paper retrieval rather than the quality of downstream literature synthesis, hypothesis generation or manuscript writing. Future work should therefore test whether improved retrieval leads to better scientific outputs in human-in-the-loop settings, including expert assessments of coverage, novelty, usefulness and trustworthiness. Third, PaSaMaster could be extended from retrieving relevant papers to helping scientists reconstruct the historical development of a field, identify emerging research trajectories and generate candidate research directions grounded in the literature. Such capabilities would help researchers rapidly understand how a topic has evolved and use verified evidence to define new research questions. Finally, although PaSaMaster eliminates source hallucination by construction, relevance scoring can still inherit biases from corpora, models and checklist design. A deployment-ready system should therefore expose evidence, uncertainty and failure modes clearly enough for researchers to audit its recommendations. Taken together, these results indicate that self-evolving, verified retrieval is a promising foundation for trustworthy AI-assisted science, but its full value will depend on transparent evaluation, continual corpus updating and careful integration into expert research workflows.
3 Methods
3.1 Overview of PaSaMaster
PaSaMaster is an agentic Recursive Self-Evolving literature retrieval system that maps a complex natural-language search intent to a ranked, evidence-grounded paper set , where is the th recommended paper and is the output list length. Its design follows three principles that directly address the limitations of existing literature retrieval paradigms: Recursive Self-Evolving retrieval, hallucination-free intent–paper relevance ranking, and cost-efficient planning–retrieval separation. Rather than generating paper lists from parametric memory, PaSaMaster retrieves real papers from customized scientific corpora, verifies their relevance using original evidence, and iteratively refines the search intent based on ranked retrieval results.
Formally, PaSaMaster operates over a customized scientific corpus and an agent-accessible operator toolset . Given query , a Navigator with policy first produces a retrieval strategy and a query-specific verification checklist , where is the number of checklist items and each checkpoint encodes one concrete requirement that a relevant paper must satisfy. The system also uses parallel Librarian agents, denoted by , where is the policy of the th Librarian. Let Plan denote Navigator planning, Retrieve denote corpus search, Verify denote evidence-grounded candidate scoring and Rerank denote final listwise ordering. The overall retrieval pipeline is:
| (1) | ||||
| (2) | ||||
| (3) | ||||
| (4) |
Here is the initially retrieved candidate set, is the evidence-scored candidate set and is the final ranked paper list. Equations 1–4 summarize PaSaMaster as a staged process in which planning defines the search, Librarians retrieve and verify papers from trusted tools, and reranking converts scored candidates into the final recommendation list.
3.1.1 Recursive Self-Evolving Retrieval from Ranked Evidence
The first core design of PaSaMaster is to transform literature retrieval from one-shot query–document matching into a Recursive Self-Evolving search process. Existing retrieval systems typically fix their interpretation of the user query at the beginning [google_scholar, asai2024openscholar, zhang2025bohriumscimaster, google2025scholarlabs, he2025pasa] and execute retrieval under this static understanding. PaSaMaster instead treats retrieval as an iterative process in which ranked evidence is used to update the system’s understanding of the research intent.
The process is coordinated by the Navigator agent. Given the initial query , the Navigator first analyzes the user’s research intent and generates two outputs: a retrieval strategy , specifying what should be searched, and a verification checklist , specifying how candidate papers should be judged. Let index the retrieval round, and let , and denote the retrieval strategy, checklist and scored candidate set at round . After each retrieval round, the Navigator inspects the ranked results, identifies missing coverage, ambiguous constraints, or under-explored directions, and refines the strategy and checklist for the next round through the reflection operator Reflect:
| (5) |
Equation 5 formalizes the self-evolving step: ranked evidence from round updates the system’s understanding of the query and produces the strategy and checklist used in round .
3.1.2 Hallucination-Free Intent–Paper Relevance Ranking
The second core design is to prevent hallucinated sources by formulating literature discovery as intent–paper relevance ranking rather than generation. PaSaMaster never asks an LLM to synthesize citations or paper lists directly from parametric memory. Instead, every candidate paper must be retrieved from a verified scientific corpus , and every relevance judgment must be grounded in traceable evidence from the original paper.
To support verifiable retrieval and evidence grounding, PaSaMaster restructures over million papers into a three-tier agent-native repository [lo2020s2orc]. The repository contains for structured metadata, for abstract-level representations used in coarse semantic filtering and for passage-level evidence chunks segmented from full texts. The corpus is represented as:
| (6) |
Equation 6 defines the corpus layout used to keep metadata retrieval, abstract-level filtering and passage-level evidence grounding separate but jointly accessible to the agents.
For each candidate paper and checklist item , let denote the set of evidence chunks belonging to paper , let denote one candidate chunk in this set, let be the shared text encoder, let denote the number of evidence chunks to retrieve, let denote selection of the highest-scoring chunks and let denote the top- evidence chunks selected to support the judgment for checklist item . The Evidence Chunk Locator retrieves supporting passages by cosine similarity:
| (7) |
Equation 7 binds each checklist judgment to explicit textual evidence by selecting the passages in a paper that are most semantically aligned with the requirement being checked.
Each candidate paper is then evaluated by a lightweight Scorer model trained through a multidisciplinary distillation pipeline. We first sample seed topics across disciplines and use them to construct initial literature queries and seed paper collections. The retrieved papers are then clustered to obtain finer-grained topic groups, from which we synthesize realistic search queries and query-specific checklists covering topical, methodological, application metadata constraints and exclusion criteria. Each synthesized query is run through the PaSaMaster retrieval pipeline to obtain candidate papers, producing multidisciplinary query–paper pairs that include positive matches, partial matches and hard negatives. A stronger teacher model then labels each pair according to the checklist, assigning checkpoint-level scores, evidence-grounded rationales and holistic relevance judgments.
This training process gives the lightweight Scorer broad, expert-style relevance assessment ability across disciplines. At inference time, for every paper and checkpoint , the Scorer outputs a satisfaction score and an evidence-grounded rationale, where larger values indicate stronger satisfaction of the checklist item. For a paper , let denote the average checklist satisfaction score over the checklist items:
| (8) |
Equation 8 summarizes checkpoint-level evidence into a single criterion-level relevance signal for candidate paper .
To incorporate holistic confidence, PaSaMaster also extracts the Scorer model’s calibrated output probability for its overall relevance judgment. Let denote the final normalized relevance score for paper . The final relevance score is:
| (9) |
where the denominator 6 normalizes the maximum possible value of , because and . Equation 9 combines checklist satisfaction and holistic confidence into one paper-level relevance score. The top candidates are then passed to a listwise reranker for global cross-paper comparison. The final result is therefore a relevance-ranked list of real papers, with each recommendation traceable to paper-level evidence.
3.1.3 Cost-Efficient Planning–Retrieval Separation
The third core design is planning–retrieval separation, which improves scalability by using frontier LLMs only where they are most valuable. Frontier LLMs are effective for understanding, decomposing, and refining complex research intents, but using them for every retrieval, reading, and ranking operation would be unnecessarily expensive. PaSaMaster therefore assigns high-level reasoning to the Navigator and delegates large-scale retrieval and relevance scoring to customized corpora, and lightweight parallel Librarian agents.
Let denote the retrieval tools used to construct candidate pools, and let denote the reading tools used to inspect metadata, abstracts and evidence chunks. The operator toolset is divided into these two subsets:
| (10) |
Equation 10 states that PaSaMaster separates tools for finding candidate papers from tools for reading and verifying them, which supports cost-efficient division of labor.
The retrieval tools construct a broad candidate pool through complementary retrieval channels, including Semantic Direct Retrieval, Citation Network Expansion, and Web-to-Repository Verification. Let denote one retrieval operator in , let denote the candidates returned by operator under strategy over corpus , and let denote the union of candidates returned by all retrieval operators:
| (11) |
Equation 11 defines the initial candidate pool as the union of complementary retrieval channels, increasing coverage before evidence verification and reranking. Semantic Direct Retrieval provides high-precision semantic candidates, Citation Network Expansion follows citation links to surface structurally related papers, and Web-to-Repository Verification maps external web findings back to verified repository entries. The reading tools then support efficient metadata lookup, abstract reading, and evidence-chunk localization, avoiding expensive full-document reading and substantially reducing computational cost.
Finally, the distilled Scorer serves as the verification component of each Librarian agent. Given a query-specific checklist and retrieved evidence chunks, it assigns checklist-level scores, generates evidence-grounded rationales and produces a holistic relevance judgment without repeatedly invoking a frontier LLM for every candidate paper. This enables Librarian agents to reproduce expert-style structured verification at much lower inference cost over large scientific corpora.
3.2 PaSaMaster-Bench
PaSaMaster-Bench is designed to evaluate literature retrieval systems under conditions that resemble real scientific paper discovery. In practice, researchers rarely ask only for a topic; they ask natural-language questions that combine scientific scope, methods, application setting, metadata restrictions and exclusions, often while searching across the open web rather than a fixed corpus. A benchmark that omits any of these dimensions can overestimate retrieval ability: short keyword-like queries understate intent reasoning, loose relevance labels do not test full constraint satisfaction, single-domain tasks hide cross-disciplinary fragility, and fixed-corpus settings bypass source verification in real search environments.
Table 3 summarizes this gap. Existing literature-search and research-agent benchmarks [he2025pasa, kang2025researcharenabenchmarkinglargelanguage, bragg2026astabenchrigorousbenchmarkingai, xiong2026autoresearchbenchbenchmarkingaiagents, ajith2024litsearch] cover useful pieces of the problem, but typically miss at least one property needed for realistic paper discovery: complex compositional natural-language intents, broad multidisciplinary coverage, or real-web search. PaSaMaster-Bench is constructed to satisfy all four properties simultaneously. It contains 244 independent tasks across 38 scientific disciplines, and each task pairs a realistic natural-language search intent with an expert-annotated target paper set .
This design makes PaSaMaster-Bench a direct testbed for the capabilities that PaSaMaster is built to provide. A system must infer the user’s complete intent, search in realistic environments, return authentic papers and filter candidates by paper-level evidence. A paper is counted as correct only if it satisfies all expert-defined checklist criteria, rather than merely sharing the same topic. The benchmark therefore evaluates whether a retrieval system can recover the intended paper set behind a real research need, not simply whether it can retrieve plausible or broadly relevant papers.
| Benchmark Dimension | RealScholar Query [he2025pasa] | Research Arena [kang2025researcharenabenchmarkinglargelanguage] | AstaBench Paper Finder [bragg2026astabenchrigorousbenchmarkingai] | AutoResearch Bench [xiong2026autoresearchbenchbenchmarkingaiagents] | LitSearch [ajith2024litsearch] | PaSaMaster- Bench |
| Natural-language queries | ||||||
| Complex search intent | ||||||
| Multidisciplinary coverage | ||||||
| Search in a real web environment |
3.2.1 Data Curation
The curation pipeline is built to preserve realism while making evaluation objective (Fig. 1b). First, domain experts write search intents grounded in authentic research scenarios across the 38 disciplines. The intent is kept in natural language so that the task resembles how a scientist would ask for papers, but it is also decomposed into a checklist of objective criteria. These criteria specify the topical scope, required methods, application setting, dataset or benchmark conditions, publication restrictions and exclusion rules that define the target paper set.
Second, each query is executed through multiple retrieval channels to approximate a real open-web search process. We use web-enabled frontier LLMs, PaSaMaster’s native search engine and traditional web search to build an intentionally broad candidate pool. The purpose is not to treat any retriever as ground truth, but to expose experts to a diverse set of possible targets, including papers that may be missed by one search channel. Retrieved papers are verified, deduplicated and mapped into a unified candidate set before annotation.
Third, domain experts annotate candidates against the checklist item by item. A paper is admitted into only when it satisfies every required criterion. Papers that are topically related but fail a method requirement, application setting, metadata constraint or exclusion rule are marked as incorrect. This elementwise annotation converts complex natural-language needs into verifiable target sets while preserving the compositional difficulty of the original query.
3.2.2 Evaluation Protocol
The evaluation protocol asks whether a system can reconstruct the target paper set implied by a complex user intent. Given a query, the system must autonomously search, verify and return a ranked list , where is the th returned paper. The list is compared with , the expert-annotated set of papers satisfying all checklist criteria. This setup makes the evaluation stricter than topical relevance: a returned paper is useful only if it satisfies the complete intent.
Coverage and ranking quality are evaluated with standard ranking metrics. At cutoff , Recall@K measures whether the system can cover the intended paper set; Precision@K measures whether returned papers satisfy the expert checklist; F1@K captures the balance between coverage and constraint satisfaction; and NDCG@K measures whether the most relevant target papers are ranked near the top [manning2008introduction, jarvelin2002cumulated, thakur2021beir]. In this benchmark, these metrics jointly test natural-language intent comprehension, paper-level filtering and ranking under compositional constraints.
Beyond retrieval quality, PaSaMaster-Bench evaluates source authenticity and cost efficiency, two requirements for deployable scientific search systems. Source hallucination rate measures the proportion of returned papers that cannot be matched to a real scholarly record or contain provably wrong source metadata. Token usage and per-query cost measure whether complex retrieval can remain low-overhead enough to scale across repeated, broad use in scientific workflows. For generative LLM baselines, we evaluate tool-assisted literature search rather than direct parametric answering: each model is equipped with Search and Visit tools and prompted to perform ReAct-style search, evidence inspection and source verification before producing a final ranked list [nakano2021webgpt, schick2023toolformer, qin2023toolllm, yao2023react, du2026openseeker, du2026openseekerv2]. This setting tests whether LLM agents can use external tools to discover real, constraint-satisfying papers rather than merely generate plausible citations.
Thus, PaSaMaster-Bench measures four capabilities required for realistic scientific literature discovery: intent comprehension, constraint-satisfying retrieval, ranking quality and source authenticity. These are also the capabilities targeted by PaSaMaster’s self-evolving retrieval, evidence-grounded ranking and planning–retrieval separation.
3.2.3 Search Intent Examples
PaSaMaster-Bench is designed around natural-language intents that mirror how researchers actually search for literature. The queries cover four common research needs in scientific work: finding a known type of study, locating target papers when only the research problem is clear, preparing a literature review with strict metadata constraints such as venue or time period, and avoiding superficially similar papers that do not meet the actual requirements. Together, these scenarios make the benchmark closer to real literature discovery than keyword-style retrieval tasks. The examples below show how PaSaMaster-Bench combines topical scope, methodological requirements, application context, metadata restrictions and exclusion criteria, requiring systems to recover papers that satisfy the complete intent rather than merely share keywords.
This is a direct intent query. The user knows what they want to find, but the target is still compositional: relevant papers must jointly concern RNA modification identification, mass spectrometry, machine-learning-based spectrum annotation and tandem-MS fragmentation evidence.
This is a problem-driven query. The user does not know the exact model family, descriptors or keywords to search for, but clearly states a research bottleneck. A capable retrieval agent must infer the relevant methodological space and translate the user’s need into search criteria before recommending papers.
This is a metadata-constrained query. The system must retrieve papers on the target topic while also enforcing publication type and time range. It therefore tests whether retrieval can combine semantic relevance with strict paper-level metadata constraints.
This exclusion-sensitive query requires selecting a specific battery system and electrolyte chemistry while rejecting nearby but incorrect targets, such as lithium-ion batteries or non-PDOL gel electrolytes. The task therefore requires judgment over both positive and negative conditions, not just semantic similarity to the topic.
Together, these examples cover much of the search behavior encountered in real scientific work: directly specified retrieval, exploratory recommendation from a research problem, metadata-restricted paper finding and fine-grained exclusion of false positives. PaSaMaster-Bench therefore evaluates both the realism and the compositional difficulty of literature discovery, requiring systems to understand what the user means, search vertically within a domain and verify whether each candidate paper truly satisfies the intended requirements.
References
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org