HAKARI-Bench:统一条件下比较检索架构与效率设置的轻量级基准
HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions
HAKARI-Bench 是一个轻量级检索基准,将现有检索套件重建为小型数据集(Nano-sets),涵盖 35 个基准、551 个任务和 43 种语言,采用统一格式实现模型无关比较。它支持 BM25、稠密、稀疏、晚交互和重排序五种检索家族及其效率变体(降维、量化等)在同一条件下对比。在 55 个模型上,整体排名与 MTEB retrieval v2、MMTEB v2 retrieval 及 English BEIR(完整版)的 Spearman 相关系数均高于 0.97。HAKARI-Bench 不取代全面评测,而是用于快速模型选择、回归检测和探索质量-效率帕累托前沿。代码、数据和排行榜以 MIT 许可证开源。
有了这个轻量级基准,做检索的开发者不用再跑整套 MTEB 就能快速筛选嵌入模型和效率配置,而且排名与完整评测高度一致,是工程选型的高性价比工具。
Abstract
With the rapid spread of retrieval-augmented generation (RAG) and semantic search, choosing the right text embedding and retrieval configuration has become both important and difficult. Large-scale retrieval benchmarks are comprehensive but too heavy to run repeatedly during development, and there is little infrastructure for comparing production-time settings—dimensionality reduction, quantization, and reranking—across many models under identical conditions. We present HAKARI-Bench, a lightweight evaluation infrastructure that reconstructs existing retrieval benchmarks into small evaluation datasets (Nano-sets) and handles benchmarks and retrieval tasks spanning languages in a unified format. Each task shares a common format of corpus, queries, relevance labels, and a fixed candidate set, enabling same-condition, model-agnostic evaluation of five retrieval families (BM25, dense, sparse, late interaction, and rerankers), together with efficiency variants: Matryoshka dimensionality reduction, int8/binary quantization, and float rescoring. Evaluating models (dense , sparse , late interaction , reranker , BM25 ), HAKARI-Bench acts as a high-fidelity ranking proxy: on the common models and intersecting tasks of each comparison, its overall ranking reproduces the official MTEB retrieval v2, MMTEB v2 retrieval, and English BEIR (full) at Spearman (, , , respectively). HAKARI-Bench is not a replacement for full evaluation; rather, it supports rapid model selection, regression detection, and reading the quality–efficiency Pareto frontier under the same conditions. The Nano-sets, evaluation pipeline, and a multi-axis leaderboard are released as open-source software under the MIT license.
Keywords: information retrieval; text embeddings; evaluation benchmark; multilingual retrieval; quantization.
Introduction
With the spread of retrieval-augmented generation and similarity search, the development of retrieval models, including text embedding models, is increasingly active. Text retrieval can be viewed as two broad stages. First is candidate generation, which retrieves candidate relevant documents from the whole corpus; alongside lexical-matching methods such as BM25, this stage uses models that represent queries and documents as dense or sparse vectors, and late interaction models that use token-level representations (Khattab and Zaharia, 2020). Second is reranking, which more precisely re-orders the top retrieved candidates; rerankers serve this stage (Nogueira and Cho, 2019). Real retrieval systems are sometimes built from candidate generation alone, and sometimes from a two-stage configuration combining candidate generation and reranking.
To compare such retrieval models, large evaluation benchmarks such as MTEB (Muennighoff et al., 2023) and MMTEB (Enevoldsen et al., 2025) have been developed, making it possible to measure performance across diverse tasks in a unified way. In information retrieval, BEIR extended to multi-domain zero-shot evaluation (Thakur et al., 2021) the test-collection format (corpus, queries, relevance labels) that TREC standardized and popularized at scale (Voorhees and Harman, 2005). BEIR organized scattered retrieval datasets into a unified format and made it possible to consistently compare the five retrieval architectures—lexical, sparse, dense, late interaction, and re-ranking—under a single model-agnostic framework (i.e., a framework in which, given the same task format and metric, any retrieval model can be swapped in for comparison).
In multilingual retrieval, MIRACL (Zhang et al., 2023) evaluates monolingual retrieval (queries and corpus in the same language) across languages, and MS MARCO (Bajaj et al., 2016), derived from Bing’s search logs, is widely used for passage retrieval. More recently, domain-specific benchmarks have proliferated rapidly: code retrieval CoIR (Li et al., 2024), long-document retrieval LongEmbed (Zhu et al., 2024), and expert-domain instruction-following retrieval IFIR (Song et al., 2025). As described later, all of these are targets that our benchmark incorporates as Nano-sets (§3.1).
On the other hand, evaluating retrieval models by measuring only overall performance on large benchmarks is not enough. In production, retrieval quality is balanced against compute, memory usage, and latency by means of dimensionality reduction or quantization of the candidate-generation embeddings, and reranking of the top candidate set. In particular, retrieval performance after output-dimension reduction based on Matryoshka representation learning (Kusupati et al., 2022), or after quantization from floating point to int8/binary (Shakir et al., 2024), is among the most-watched areas after the base model performance itself. Hence, in addition to base model performance, it is important to understand how performance changes when these efficiency settings are used, or when a two-stage configuration of candidate generation followed by reranking over its candidate set is adopted.
However, existing large benchmarks cover many tasks at a large data scale, so re-evaluating all other models under the same conditions after changing dimensionality reduction or quantization for one model is not easy. In particular, the scaled-up MMTEB (Enevoldsen et al., 2025) infrastructure can in principle handle dimensionality reduction based on Matryoshka representations (Kusupati et al., 2022) and quantized embeddings (Shakir et al., 2024), but in practice these efficiency settings are rarely reported consistently per model, and there are almost no results comparing multiple models under the same conditions. Moreover, frameworks that systematically compare, under the same conditions and according to their respective roles, the architecture that generates candidates from the whole corpus and the architecture that re-orders top candidate sets, are limited. From the same motivation, in the English domain NanoBEIR lightweights each BEIR dataset and is widely used as a fixed-dataset ranking proxy (Câmara, 2024; Aarsen, 2024).
In this paper we build HAKARI-Bench, a lightweight benchmark for evaluating multilingual, multi-domain retrieval models.111The evaluation and visualization implementation is open source under the MIT license: https://github.com/hakari-bench/hakari-bench. The leaderboard of evaluated models is public at https://huggingface.co/spaces/hakari-bench/leaderboard; the evaluation data (Nano-sets) is released on Hugging Face Datasets (Appendix G). The name “HAKARI” comes from the Japanese word for a weighing scale (hakari, “to measure”), reflecting the benchmark’s aim of measuring and comparing retrieval models.
Specifically, we construct small evaluation datasets (hereafter Nano-sets) from existing retrieval benchmarks and develop an infrastructure that handles benchmarks and retrieval tasks uniformly. Each task is handled in a common format consisting of a corpus, queries, relevance labels, and a top candidate set, so that candidate-generation methods and reranking methods can be evaluated with the same metrics according to their respective roles. We further make it possible to evaluate, under the same conditions, candidate-generation methods such as BM25, dense, sparse, and late interaction, as well as reranker evaluation over the top candidate set, embedding dimensionality reduction, int8 quantization, and binary quantization.
In this paper, “lightweight” refers solely to reduced evaluation cost (the ease of repeated measurement enabled by Nano-set construction). Dimensionality reduction, quantization, and sparse pruning are evaluated as reproducible proxies for storage and retrieval cost (embedding dimension, quantization precision, number of non-zero dimensions); we do not evaluate the inference speed itself of each model, because fair measurement is difficult (§7.5).
The positioning of our benchmark is summarized in three points. First, it follows the consistent evaluation methodology established by BEIR (Thakur et al., 2021), i.e., same-condition comparison of multiple retrieval architectures based on a unified format. Second, it extends this to many models and to many languages and domains. Third, by shrinking the data size of each task, it makes retrieval tasks repeatedly measurable at a realistic speed, and on top of that applies dimensionality reduction, int8 quantization, and binary quantization to all supporting models under the same conditions. This means providing, in a consistent manner across all target models, the comparison of efficiency settings that is possible on the evaluation infrastructure but has in fact been measured only sporadically per model.
We also verify the extent to which Nano-sets reproduce the model ranking of the original large-scale evaluation. Comparing HAKARI-Bench’s Nano-set results against MTEB retrieval v2, MMTEB v2 retrieval, and English BEIR (full), the Spearman rank correlations were , , and , and the Pearson correlations were , , and , respectively. In addition to rank correlation itself, high correlation was obtained for the Borda score that aggregates per-task wins/losses, indicating that while HAKARI-Bench does not replace full evaluation, it reproduces model ranking with high fidelity and functions as a lightweight evaluation metric.
Given the established fact that neural retrieval models degrade substantially out of distribution (Thakur et al., 2021), the true value of a multi-domain benchmark is not maximizing the overall score, but exposing “which domains a model has not learned” and providing material for use-appropriate model selection. Our benchmark is designed so that tasks can be sliced and compared along axes such as language, domain, and query length, supporting this perspective (§5.7, §6.2).
Based on the above, the contributions of this paper are threefold.
- 1.
A lightweight multilingual, multi-domain retrieval evaluation infrastructure. We reconstruct existing retrieval benchmarks as Nano-sets and build an infrastructure that compares the five families of BM25, dense, sparse, late interaction, and reranker in a unified format under the same conditions over benchmarks and tasks (§3, §4).
- 2.
Empirical validation of ranking reproducibility of the lightweight evaluation. We show, through three independent comparisons, that the overall ranking induced by Nano-sets reproduces the official MTEB retrieval v2 / MMTEB v2 retrieval and BEIR (full) at Spearman in every case, on the common models and intersecting tasks of each comparison (§5.6, §6.1).
- 3.
Cross-model evaluation of efficiency settings and reranking. We applied dimensionality reduction, int8/binary quantization, and rescoring to all supporting models under the same conditions, and evaluated reranking over a fixed candidate set on all tasks. These settings are in principle measurable on existing infrastructure, but in practice have been reported only sporadically per model. Our contribution is not opening a new measurability for the first time, but actually providing it, by applying it uniformly to all supporting models so that efficiency and reranking performance can be compared across models on the same basis. This makes concretely visible the differences that emerge only when settings are held fixed—for example, that robustness to binary quantization is determined by a model’s training characteristics (not explained by size or dimension), and that whether a reranker beats dense changes with the task type and architecture (§5.3–§5.5, Appendix F).
Related Work
Retrieval evaluation benchmarks and retrieval architectures
The evaluation format for text embedding models was standardized when MTEB (Muennighoff et al., 2023) unified eight tasks including retrieval, reranking, classification, and clustering. MMTEB (Enevoldsen et al., 2025) extended the scope to over languages and over tasks, and introduced quality review and correlation-based downsampling. Restricting attention to information retrieval, BEIR (Thakur et al., 2021) assembled a zero-shot IR suite of datasets and made it possible to consistently evaluate the five retrieval architectures—lexical, sparse, dense, late interaction, and re-ranking—under a single model-agnostic framework. This design of “comparing different retrieval architectures under the same conditions” is the direct origin of our evaluation methodology (§4).
For multilingual monolingual retrieval, MIRACL (Zhang et al., 2023) is widely used; for short-query passage retrieval, MS MARCO (Bajaj et al., 2016); and for domain specialization, code retrieval CoIR (Li et al., 2024), long-document retrieval LongEmbed (Zhu et al., 2024), instruction-following retrieval FollowIR (Weller et al., 2024), and expert-domain IFIR (Song et al., 2025) have been developed.
More recently, the official MTEB leaderboard introduced a retrieval-specific section centered on RTEB (Retrieval Embedding Benchmark; Liu et al., 2025). RTEB is a retrieval-focused benchmark that measures multilingual retrieval quality across production domains such as legal, finance, code, and medical; it combines public datasets with private (closed) datasets to be robust to training-data contamination and leaderboard overfitting. However, like MTEB, RTEB evaluates embedding models under a fixed protocol and does not primarily aim at cross-architecture comparison (lexical, dense, sparse, late interaction, re-ranking) or comparison of efficiency settings such as dimensionality reduction and quantization. Our HAKARI-Bench moves in step with this retrieval-focused trend, but is complementary in that it performs cross-architecture comparison and efficiency-setting evaluation on top of lightweight measurement via Nano-sets.
The retrieval architectures under evaluation divide broadly into two stages: candidate generation, which retrieves candidates from the whole corpus, and reranking, which re-orders top candidates. Candidate generation commonly uses lexical-matching BM25 (a robust baseline even for zero-shot IR; Robertson and Zaragoza, 2009; Thakur et al., 2021), dense retrieval with bi-encoders (Reimers and Gurevych, 2019; Karpukhin et al., 2020), learned sparse SPLADE-family models (Formal et al., 2021), and token-level late interaction (ColBERT family; Khattab and Zaharia, 2020; Santhanam et al., 2021). Reranking, starting from two-stage retrieval that re-orders BM25 candidates with a BERT cross-encoder (Nogueira and Cho, 2019), is a re-ordering of top retrieved candidates, not a retrieval over the whole corpus. Because the two have different roles, they should be evaluated according to their respective roles rather than compared in the same role.
These benchmarks have greatly contributed to the comprehensive comparison of retrieval models, but the more comprehensive and large-scale they become, the harder it is to repeat the full evaluation during development. We restrict our scope to retrieval and reranking precisely because a lightweight, same-condition infrastructure is needed to iteratively check retrieval-specific comparison axes such as candidate generation, reranking, and efficiency settings.
Lightweight evaluation and Nano-set construction
To lower the cost of repeated evaluation on large benchmarks, the evaluation data has been lightweighted. MMTEB (Enevoldsen et al., 2025) shrinks tasks through correlation-based downsampling, showing a policy for obtaining conclusions close to the full evaluation at low cost. As a smaller-scale lightweighting, NanoBEIR is a collection that shrinks each BEIR dataset to about queries up to K documents; it was introduced by Zeta Alpha for evaluation-cost reduction (Câmara, 2024) and unified into a single format by Sentence Transformers as the NanoBEIREvaluator (Aarsen, 2024). Negative documents are sampled with Pyserini’s BM25 and a general-purpose dense model. For lightweight reranker evaluation, a derived collection with BM25 candidate scores attached to the Nano-sets has also been prepared (Sentence Transformers, 2024). Multilingual extensions have been released as translated/improved versions by LightOn AI (Sourty, 2025), Liquid AI (Liquid AI, 2025), and Sionic AI (Sionic AI, 2025).
Our Nano-set construction follows this idea of a “ranking proxy on a small collection” and the practice of fixing BM25 candidates, and is distinctive in extending the net to non-English languages, expert domains, and comparison of efficiency settings including dimensionality reduction and quantization.
Evaluating embedding efficiency settings
In production, embedding dimensionality reduction and quantization are widely used to balance retrieval quality against compute, memory, and latency. Matryoshka representation learning (Kusupati et al., 2022) is a method that trains embeddings so that the dimensions can be truncated while preserving the leading dimensions, giving an axis for comparing retrieval performance after dimensionality reduction. Embedding Quantization (Shakir et al., 2024) combined int8/binary quantization with float rescoring to reduce storage and retrieval cost. In particular, the two-stage configuration of “generating candidates efficiently with binary codes and re-ranking (rescoring) accurately with continuous vectors” traces back to the Binary Passage Retriever (Yamada et al., 2021). Production approximate nearest neighbor (ANN) search uses more advanced quantization, such as Product Quantization (Jégou et al., 2011), Optimized PQ (Ge et al., 2013), RaBitQ with a theoretical error bound (Gao and Long, 2024), Better Binary Quantization with correction terms (Trent, 2024), Optimized Scalar Quantization (Veasey, 2026), and TurboQuant, which combines Hadamard rotation with re-normalization and calibration (Pijpelink, 2026).
However, evaluation infrastructure that can compare the impact of these efficiency settings on retrieval quality across many models under the same conditions is limited. Large-scale infrastructure such as MTEB / MMTEB (Muennighoff et al., 2023; Enevoldsen et al., 2025) can in principle handle dimensionality reduction, but performance changes after quantization or dimensionality reduction are often confined to per-model initialization settings, and cross-model same-condition comparisons are rarely reported. We treat these efficiency settings as first-class records of the evaluation results and compare quality and efficiency side by side in the same table (§4.3, §5.3).
Positioning relative to existing benchmarks
We summarize the relationship between the existing benchmarks discussed above and HAKARI-Bench in Table 1. In the table, = consistently provided as a first-class feature across retrieval tasks, = limited (possible on the infrastructure but not consistently reported across models, or restricted to dedicated tasks / separately distributed data), = out of scope.
| Aspect | BEIR | MTEB | MMTEB | NanoBEIR | HAKARI-Bench |
|---|---|---|---|---|---|
| Retrieval task scale | 18 datasets | retr. 15 | 500+ tasks | 13 datasets | 551 tasks / 35 bench. |
| Languages | English-centric | English-centric | 250+ langs | English (multiling. derivs) | multilingual (43) |
| Consistent cross-architecture | 5 fam. | dense | dense | dense+BM25 | 5 fam. |
| Reranker eval. (on task cand.) | ad-hoc BM25 | dedicated only | dedicated only | separate BM25 set | fixed set, all tasks |
| Dim. reduction (Matryoshka) | |||||
| Quantization (int8/binary) | |||||
| Leaderboard | (dataset set) | (multi-axis) | |||
| Repeated-measurement cost | heavy | medium–heavy | medium (downsamp.) | light | light |
Regarding reranker evaluation: BEIR’s re-ranking is described as re-ranking the top first-stage BM25 hits (Thakur et al., 2021), and is not, as in this paper, a design that distributes a single fixed candidate set to all models for rescoring. MTEB / MMTEB reranking consists of a few dedicated tasks and is not a re-ordering over the candidate set of the retrieval task itself (Muennighoff et al., 2023; Enevoldsen et al., 2025), and NanoBEIR requires a separately distributed BM25-candidate-augmented derived collection (Sentence Transformers, 2024). By contrast, HAKARI-Bench ships a fixed hybrid candidate set for all retrieval tasks and evaluates rerankers consistently on the same candidate set as candidate generation (§3.3, §4.2). On repeated-measurement cost, MTEB retrieval originally uses BEIR’s full corpora (hundreds of thousands to millions of documents) and is heavy; MTEB v2 retrieval partly adopts MMTEB-derived hard-negative downsampling to shrink the corpus (the v2 version is what we compare against in §5.6), and is moderately lightweighted within that scope.
Overall, BEIR established the evaluation design of consistent cross-architecture comparison, but is English and full-scale, with cross-model evaluation of efficiency settings out of scope. MTEB / MMTEB extended to multilingual, multi-task settings and reduced measurement cost via downsampling (Enevoldsen et al., 2025), but dimensionality reduction and quantization, though handleable on the infrastructure, are not reported consistently per model, and reranking is limited to dedicated tasks. NanoBEIR achieved lightweighting in the English domain (Câmara, 2024; Aarsen, 2024), but efficiency settings are out of scope. HAKARI-Bench inherits these strengths—BEIR’s consistent methodology, MMTEB’s multilinguality, NanoBEIR’s lightness—while integrating reranker evaluation on the retrieval task candidate set and cross-model evaluation of efficiency settings into a single evaluation infrastructure.
Design of HAKARI-Bench
HAKARI-Bench is not merely a dataset collection but an evaluation infrastructure that handles the task set, candidate-generation evaluation, reranking evaluation, and efficiency settings as a whole. It is a five-stage pipeline.
- 1.
Task specification. The dataset location and version (commit SHA), language, and domain category are written as a declarative configuration file (§3.2).
- 2.
Common task format. Each task is aligned to a corpus, queries, relevance labels (qrels), and a fixed top candidate set (by default the hybrid top obtained by fusing BM25 and dense with RRF) (§3.1, §3.3).
- 3.
Evaluation. The five families of BM25, dense, sparse, late interaction, and reranker, together with efficiency variants (dimensionality reduction, int8, binary, rescore, sparse pruning), are run on the same tasks (§4).
- 4.
Result records. All runs are stored in a single schema (per-query top ranking, various @k scores, variants, resolved versions, diagnostic records).
- 5.
Aggregation and display. Results are aggregated into a DuckDB warehouse and displayed as a leaderboard with macro/micro averages and multi-axis filters (§4.5).
This section describes that design.
Task set and Nano-sets
The basic unit of the benchmark is a retrieval task. Each task adopts the test-collection format that TREC standardized and popularized at scale (corpus, queries, relevance labels qrels; Voorhees and Harman, 2005), augmented with a fixed top candidate set (by default the hybrid top fusing BM25 and dense, §3.3). Just as BEIR unified scattered IR datasets into this format (Thakur et al., 2021), our benchmark adopts the same format as a common interface and consistently advances multilingual, expert-domain, and Nano-set development.
Each task is constructed as a small evaluation dataset (Nano-set) shrunk from the original benchmark to about – queries and about K–K documents. This shrinking, inspired by MMTEB’s downsampling (Enevoldsen et al., 2025) and NanoBEIR’s Nano-sets (Câmara, 2024; Aarsen, 2024), aims to lower the cost of repeated evaluation.
Nano-set construction is twofold by provenance. First, already-published Nano collections such as the NanoBEIR family are referenced by name and version on the Hugging Face Hub without re-implementing the individual shrinking logic. Second, families that we reconstruct from the official MTEB / MMTEB full evaluation (NanoMTEB-v2, NanoMMTEB-v2, etc.) are made into Nano-sets by a common shrinking procedure: (i) select up to deduplicated queries that have at least one positive qrel, and (ii) for the corpus, after including all positive documents of the selected queries, cap it at about K documents, preferentially adding any hard negatives present in the original data (documents explicitly labeled non-relevant, i.e., qrels with score 0) in a query-crossing round-robin, and filling the remainder with documents in the original corpus order.
For tasks whose original data has no hard negatives, the filler documents make up most of the candidate space, so the retrieval space becomes relatively easy (irrelevant documents are unlikely to be incidental hard negatives), making it easier to distinguish positives from queries. This construction can be applied uniformly to many tasks at low cost, but there is room to raise the discriminative power of Nano-sets, e.g., by adding hard negatives (§8). Even so, the NanoMTEB-v2 / NanoMMTEB-v2 reconstructed this way retain a sufficient rank correlation with the official evaluation, as confirmed in §5.6. Storing a fixed BM25 top- candidate set on the dataset side follows the same idea as Sentence Transformers’ BM25-candidate-augmented derived collection (Sentence Transformers, 2024), decoupling reranker and learned-sparse evaluation from per-run BM25 computation differences.
The benchmark contains benchmarks and retrieval tasks. Task selection follows the four criteria BEIR identified (task diversity, domain diversity, task difficulty, coexistence of annotation strategies; Thakur et al., 2021), extended to multilingual and expert domains. The task set comprises the following five families. This taxonomy is a convenient organization based on benchmark provenance and target, not a distinction in the evaluation implementation. We give the main source benchmark for each Nano-set here; details of version, provenance, and number of languages are organized in Appendix A.1 (Table LABEL:tab:a1).
- •
BEIR family: MNanoBEIR, a multilingual collection integrating the English version that Sentence Transformers reformatted (Aarsen, 2024) from Zeta Alpha’s original NanoBEIR collection (Câmara, 2024) with the translated/extended multilingual derivatives by LightOn AI (Sourty, 2025), Liquid AI (Liquid AI, 2025), and Sionic AI (Sionic AI, 2025) (original datasets are BEIR; Thakur et al., 2021), comprising BEIR datasets language editions. At aggregation time it is grouped hierarchically by language and dataset and treated as one benchmark like the others (§4.5).
- •
Official MTEB family: aligned with the official MTEB / MMTEB v2 (Muennighoff et al., 2023; Enevoldsen et al., 2025) and separated per official family: NanoMTEB-v2, NanoMMTEB-v2, NanoCMTEB, NanoJMTEB-v2, NanoFaMTEB-v2, NanoRuMTEB, NanoVNMTEB, NanoMTEB-Misc, and per-language NanoMTEB-{Dutch, French, German, Korean, Polish, Scandinavian, Spanish, Thai}. The per-language source benchmarks each family references (C-MTEB, MTEB-NL, MTEB-French, SEB, ruMTEB, VN-MTEB, etc.) are given in Table LABEL:tab:a1.
- •
Multilingual general: NanoMIRACL (Zhang et al., 2023), NanoMLDR (Chen et al., 2024), NanoIndicQA (Doddapaneni et al., 2023), NanoMuPLeR (built by MTEB from the EU DGT multilingual parallel corpus; Table LABEL:tab:a1). Each spans multiple languages.
- •
Long-document, instruction-following, expert-domain, reasoning: NanoLongEmbed (Zhu et al., 2024), NanoIFIR (Song et al., 2025), NanoChemTEB (Shiraee Kasmaee et al., 2024), NanoR2MED (Zhang et al., 2025a, R2MED), NanoBIRCO (Wang et al., 2024b), NanoBRIGHT (Su et al., 2024), NanoRARb (Xiao et al., 2024a), NanoRTEB (RTEB; Liu et al., 2025, here English production-domain retrieval (legal, finance, code, etc.), not multilingual), NanoBuiltBench (BuiltBench; Table LABEL:tab:a1), NanoDAPFAM (DAPFAM; Table LABEL:tab:a1), and the composite tasks NanoLaw (legal IR composite; AILA, LegalBench, etc., Table LABEL:tab:a1) and NanoMedical (medical IR composite; CURE, etc., Table LABEL:tab:a1).
- •
Code: NanoCoIR (Li et al., 2024), NanoCodeRAG (Wang et al., 2025).
The task set contains duplicate tasks that derive from the same original dataset across families (e.g., scidocs, trec_covid). Because the Nano-set sampling differs by family, the same original task can become a different evaluation surface, so we keep duplicates as independent tasks rather than merging or removing them. Duplicate tasks may be double-counted in the equal-weight micro average over all tasks, but this effect is mitigated in the per-benchmark macro average that our analysis uses as the primary basis (cross-benchmark micro/macro and the default display are discussed in §4.5). The list and details of duplicates are in Appendix A.3.
The overall picture of benchmarks/tasks and the distribution of languages and document counts are shown in §5.1; the provenance and version of each Nano-set are organized in Appendix A.
Common evaluation format
All tasks are aligned to a common format of corpus, queries, relevance labels, and top candidate set. This unification makes it possible to compare candidate-generation methods (BM25, dense, sparse, late interaction; retrieving the top from the whole corpus) and reranking methods (re-ordering the candidate set) with the same metrics according to their respective roles (§4). Evaluation results are stored in a single schema so that swapping models, evaluation methods, prompts, and efficiency variants can be handled by a common pipeline; task specifications are managed as declarative configuration files whose required metadata are the dataset location and version, language, domain category, and citation information (a collection of task specifications becomes one benchmark on the leaderboard).
Top candidate set
Reranking is not whole-corpus retrieval but a re-ordering over a top candidate set. To make this premise explicit, we fix and share the candidate set per task. The candidate set defaults to a hybrid candidate set (top ) fusing the BM25 top and the dense-retrieval top via RRF (Reciprocal Rank Fusion), and the reranker and the candidate-generation baselines share the same candidate set. The construction is as follows. For each query we retrieve the top from the whole corpus with BM25 and the top with a fixed dense model (microsoft/harrier-oss-v1-270m, using the dedicated prompt web_search_query), then fuse the two rankings with RRF (each document scored by , ) and take the top . We use a fixed dense model for dense retrieval to decouple candidate-set construction from the models under evaluation and to fix a reproducible candidate pool consistently across all tasks. The dataset side also stores a BM25-only candidate set (top ) as a lexical baseline, switchable as needed. Fixing the candidate set on the dataset side makes a reranker’s improvement less dependent on candidate-generation bias. In particular, because the hybrid candidate set contains not only BM25 but also the dense top, it is a shared re-ordering target that is not skewed to a single candidate-generation method (BM25-only or dense-only); it nonetheless depends on the specific hybrid construction (the BM25/dense tops, RRF, and the positive-append safeguard), as discussed in §7.3.
This fixed candidate set has a safeguard rule: only when the top contains no positive at all do we append one positive at the tail (rank ), ensuring that every query has at least one relevant document in the candidate set (query coverage ). Since passing a candidate set with no positives to a reranker yields no meaningful evaluation signal, this is a design decision that prioritizes isolating reranker evaluation to “ranking accuracy over the candidates.” Inclusion of all relevant documents (relevant-document coverage) is not guaranteed (about on dense average; §5.5), and candidate-generation failures are observed as an axis independent of reranking evaluation (§5.5, Appendix E.4). This is also due to the exceptional situation that, in addition to candidate generation missing some positives, for tasks with many positive documents per query it is in principle impossible to include all of them in a capped -document candidate set (the safeguard adds only one). The implications of this design, including its difference from real-world two-stage retrieval, are discussed in §7.3. Also, to treat BM25 fairly across languages in multilingual monolingual retrieval, the BM25 computation for the candidate set uses per-language tokenizers (morphological analyzers for CJK, Thai, and Vietnamese; Unicode regular expressions plus stemming for some languages otherwise). Details of the candidate-set construction, including the without-safeguard metric and the tokenizer breakdown, are in Appendix E.3.
Evaluation targets
The benchmark evaluates candidate generation, reranking, dimensionality reduction, quantization, and sparse-representation pruning on the same task set. Candidate-generation methods (BM25, dense, sparse, late interaction) are evaluated as methods that retrieve candidates from the whole corpus. Rerankers are evaluated as re-ordering over the top candidate set. For dense embeddings, we derive, as variants, leading-dimension-preserving dimensionality reduction, int8 quantization, binary quantization, and their combinations, comparing quality and efficiency side by side. For sparse representations, we evaluate how far the query-side and document-side representations can each be pruned. In evaluation, the dataset version (commit SHA) can be specified explicitly, and the resolved SHA is recorded in the results, so correspondence with past numbers is preserved even when a dataset is updated (Appendix A.2).
Evaluation Methodology
This section describes which models are evaluated as candidate generation, which as reranking, and with what metrics. The evaluation modes the benchmark provides align with the five retrieval architectures BEIR identified (lexical / sparse / dense / late interaction / re-ranking; Thakur et al., 2021); every mode takes the same task specification as input and outputs the same result schema.
Evaluating retrieval models
Retrieval models are evaluated as methods that retrieve candidate relevant documents from the whole corpus. BM25 is a lexical-matching method; by default the stored BM25 top is evaluated (local computation is also switchable). Dense retrieval encodes queries and documents with an embedding model and retrieves the top from exact similarity over the whole corpus. For each model–task pair we compute both cosine and inner-product similarity and report whichever yields the higher task nDCG@10; this is a per-task best-of-similarity upper bound over the two functions (an oracle over the similarity choice), applied uniformly to all dense models (Appendix C.3) (dimensionality-reduction and quantization variants are in §4.3). Sparse retrieval scores by the inner product of learned sparse representations (pruning settings in §4.4); late interaction scores by the token-to-token MaxSim of ColBERT-family token-level embeddings. In every method, the top retrieval results per query can be stored, so downstream reranking and error analysis can be run without recomputing embeddings.
Evaluating rerankers
A reranker is evaluated not by comparison in the same role as candidate generation, but as a re-ordering over the top candidate set. In this paper, a reranker is a model that takes a query–document pair as input, directly scores their relevance, and re-orders the candidate set. The representative example is a BERT/XLM-R cross-encoder (Nogueira and Cho, 2019), but we also include LLM-style rerankers based on a large language model (decoder) that use the predicted logit of the “yes / no” token as the relevance score. We collectively call these rerankers, whether cross-encoder or LLM-style. A reranker re-orders the fixed candidate set (by default the hybrid top ) and we compute post-reranking metrics. Because the candidate set contains at least one relevant document for every query under the safeguard rule (§3.3; query coverage ), reranker evaluation focuses on ranking accuracy over the candidates.
Not only dedicated rerankers but also retrieval models (dense, sparse, late interaction) can be scored as rerankers by rescoring the same fixed candidate set. Hence reranking performance can be measured under the same conditions for both dedicated rerankers and retrieval models. In particular, the improvement when a retrieval model re-evaluates its own candidate set and the performance when a reranker re-orders the candidate set can be read separately on the same candidate set (§5.5).
Dimensionality reduction and quantization of dense embeddings
Efficiency settings for dense embeddings are generated as derived variants by post-encoding transformations after computing the base embedding once. This derives multiple efficiency settings from a single inference under the same conditions and compares quality and efficiency. The variants are:
- 1.
Dimensionality reduction (truncation): a leading-dimension-preserving dimension slice (assuming the Matryoshka family; Kusupati et al., 2022). We compare performance when truncated to, e.g., dimensions.
- 2.
Quantization: int8 and binary. int8 is not a type cast to float16 but a scalar quantization that linearly quantizes each dimension to an 8-bit integer ( levels). Concretely, we take per-dimension min/max from the corpus-side embeddings and map each value into one of buckets spanning that range (per-dimension affine quantization). Calibration is done only on the distribution-stable corpus side; queries are not used for calibration (to avoid fitting buckets to evaluation queries), and out-of-range values are clipped. No separate calibration sampling or training is performed (same family as the quantization of Shakir et al., 2024). Binary keeps only the sign of each dimension (1-bit).
- 3.
rescore: the simplest two-stage retrieval, which rescores the top retrieved by quantized search using the original floating-point embeddings.
- 4.
Combinations: the cross product of dimensionality reduction quantization rescore.
Each variant is stored side by side as a separate record for the same task, and the leaderboard’s “delta vs. base” column directly shows quality degradation. This lets the leaderboard be read not as a single score column but as a Pareto frontier of quality and efficiency. Here a Pareto frontier is the set of settings on the two axes of quality and efficiency (embedding dimension, quantization precision, etc.) that cannot be beaten without worsening one of the two; i.e., the locus of best quality reachable for a given efficiency, and conversely the locus of minimum cost reachable while preserving a given quality.
Note that our quantization is a simple post-hoc quantization for measuring a model’s own quantization robustness; the gap to advanced production ANN methods (§2.3) is discussed in §6.4. Technical details of the variants are organized in Appendix E.1.
Sparse-representation pruning settings
Because learned sparse representations are inherently sparse, it is common to keep only the top-absolute-value dimensions per row (a max active dims limit). We measure, on the same evaluation surface, single variants that independently specify the query-side and document-side max active dims, plus their combination variants. The query-side value determines the number of non-zero dimensions at search time and is directly tied to search latency. The document-side value, in addition to latency, is directly tied to the size of the inverted index and embedding matrix, i.e., the production-time memory/disk footprint. Listing the two independently lets us read the relationship between pruning settings and retrieval quality for a given operating environment (latency budget, memory/storage budget) (§5.4, Appendix E.2).
Metrics and aggregation
The benchmark’s main metric is nDCG@10, following the primary metric BEIR adopted (Thakur et al., 2021). A key design point is that during each task’s evaluation, the per-query top ranking is stored as an artifact. With the top ranking stored, various retrieval metrics (nDCG, recall, accuracy, MRR, MAP, etc.) can be recomputed at any time from the stored rankings when building the leaderboard (DuckDB warehouse). The viewer/leaderboard default display is the main metric nDCG@10, recorded in ; co-reporting recall@ as a secondary metric follows BEIR’s convention. Metric definitions are detailed in Appendix B.1.
To robustly aggregate benchmark groups with skewed task counts and scales, we follow these rules. Per benchmark, we display the simple average over tasks (). For cross-benchmark aggregation, we co-report the equal-weight micro average over all tasks and the macro average that equally weights each benchmark. The leaderboard/viewer default display is the micro average, with macro equally switchable. We default to micro because, combined with language/category filters or Nano-set narrowing, “equal-weight average over all tasks in the displayed range” is a simple, easy-to-understand interpretation. Which aggregation is appropriate depends on “what data, at what granularity, one wants to see,” so the two are placed side by side and switchable. In our analysis, however, to avoid the overall score being dominated by benchmark groups with skewed task counts and scales (especially the -task MNanoBEIR and cross-family duplicate tasks), we report the per-benchmark macro average as the primary aggregation basis. In macro aggregation, the BEIR-family MNanoBEIR ( BEIR datasets languages) is first averaged over the languages within each BEIR dataset, and the dataset averages are then averaged into a single benchmark score (hierarchical aggregation by language and dataset). This prevents the high-row-count MNanoBEIR from dominating aggregation by task-count weight. Ranking targets only models that have the entire expected task set within the selected display range. Aggregation details and handling of missing tasks are in Appendix B.2, B.3.
Results can be displayed through multi-axis filters based on task and model metadata. Representative axes are (i) language tags (e.g., comparing a Japanese-specialized model side by side with multilingual general models), (ii) domain category (code/natural language, expert domain), (iii) per-task average query length and average document length (e.g., excluding long-document tasks when a model not trained on long context produces an extremely low score that distorts the overall ranking), and (iv) model embedding dimension and parameter count. These are means of reconstructing the leaderboard along use-appropriate cuts, supporting the separation of a model’s strong and weak domains that is hard to see in a single score over the whole task set (§6.2).
Results
The evaluation values reported below are based on a fixed HAKARI-Bench snapshot as of 2026-06-09 (DuckDB warehouse hakari-bench/leaderboard_database commit 1f0d59d, build 2026-06-09, schema v8); the official mteb/results used for the rank-correlation comparison (§5.6, Appendix D) is fixed at commit 1e8ab5d, as of 2026-06-08. These two snapshots are the fixed reference basis of the paper. Hereafter “the present snapshot” refers to this data snapshot (build 2026-06-09) and is used without further qualification.
The most important result of this section, stated up front: the overall ranking induced by Nano-sets reproduces the official full evaluations (MTEB retrieval v2 / MMTEB v2 retrieval and BEIR) at Spearman (§5.6), confirming that the lightweighting does not damage the ranking proxy. The overall values in this section use the per-benchmark macro average as the primary basis to suppress task-count skew (the leaderboard/viewer default display is the micro average; when the difference between the two affects interpretation we note it explicitly; §4.5). Below we first show the evaluation targets and task distribution (§5.1), then model performance and efficiency-setting results (§5.2–§5.5), the rank correlation that grounds their validity (§5.6), and real-data use cases (§5.7).
The evaluation targets include base rows of models222The fixed DuckDB snapshot itself contains models ( dense); we exclude two unreleased dense models from all pools, aggregations, and figures/tables, leaving the ( dense) analyzed here (Appendix C.1). benchmarks tasks. The models comprise, as candidate-generation methods, dense embeddings, learned sparse, late interaction (ColBERT family), and lexical-baseline BM25, plus rerankers ( cross-encoders, LLM-style) that re-order the top candidate set. Except for BM25, all are small-to-medium models of about B parameters or fewer; the benchmark mainly targets the band that distributes at about B or fewer on the MMTEB leaderboard. This lets all five families BEIR defined (lexical / sparse / dense / late interaction / re-ranking) be compared on the same task set. The model composition and references are in Appendix C.1.
Task-set distribution
The benchmark is not a lightweight version of a single domain but a multilingual, multi-domain evaluation surface. The category distribution is natural-language tasks ( queries, documents) and code tasks ( queries, documents). Five benchmarks contain code tasks, of which NanoCoIR and NanoCodeRAG are code-only and NanoBRIGHT, NanoRTEB, and NanoRARb are mixed with natural-language tasks (per-benchmark task counts are in Table LABEL:tab:a1). The languages tagged in the task metadata span languages in total. The top languages by per-language task count (counting a task under each of its languages when it spans multiple) are English , Vietnamese , German , French , Dutch , Japanese , Spanish , Thai , Korean , and Arabic/Persian each; the cumulative number of tasks for non-English languages exceeds . Note these are task counts, not language counts; the number of distinct target languages is . All tasks have complete metadata for query count, document count, and average character length.
Overview of model performance
The per-benchmark task average (, on a -dense-model basis) varies greatly across benchmarks. High benchmarks include NanoCodeRAG , NanoRuMTEB , NanoChemTEB , NanoCoIR , and NanoMIRACL ; low benchmarks include NanoRARb , NanoR2MED , NanoDAPFAM , NanoBIRCO , and NanoBRIGHT . Even on the Nano-set task collection, differences of over points are observed across benchmarks. This reflects, rather than wins/losses of individual models, the exposure of existing embedding models’ weaknesses on expert-domain, instruction-following, and complex-reasoning tasks, and on natural-language tasks a model does not support (e.g., English-only models degrade greatly on multilingual tasks): performance differences vary greatly by task, domain, and supported language (§6.2).
Looking at the overall ranking of dense models by per-benchmark macro average (), the top are jinaai/jina-embeddings-v5-text-small (jina-embeddings-v5; Akram et al., 2026) , jinaai/jina-embeddings-v5-text-nano , microsoft/harrier-oss-v1-0.6b , perplexity-ai/pplx-embed-v1-0.6b , and google/embeddinggemma-300m (the equal-weight micro average gives , , , , , respectively, with a stable top composition). Because the macro average is less pulled by large benchmarks, we use it as the primary overall basis (the leaderboard default display is micro, with macro switchable; §4.5). The lexical-baseline BM25 scores macro (micro ) under full-corpus retrieval; though below the top dense group, it is co-reported on all tasks as the baseline for same-condition comparison across architectures. Fine distinctions between nearby models cannot be settled by a single Nano-set with limited queries, and should be read as a ranking proxy (§7.2).
Performance change from dimensionality reduction and quantization
int8, binary, and their rescore variants are complete over dense models tasks. Matching these variants against the base rows on the same tasks and taking the all-model mean of the delta vs. base (, i.e., points), binary is points, int8 , binary_rescore , and int8_rescore . That is, binary quantization alone has the largest quality drop, int8 is mild, and adding rescore restores int8 to almost lossless () and binary to . Here, rescore means rescoring the top retrieved by the quantized vectors using the original floating-point (e.g., fp16) embeddings retained before quantization, then re-ordering (§4.3). Production search engines often retain the original vector values; then the rescoring targets only a few top candidates, so the additional compute is small (one may even recompute with the original model’s non-reduced, non-quantized vectors). The recovery by rescore shows that quality is almost entirely regained at this small cost. This is a trend that can only be confirmed cross-model by applying the efficiency settings to all supporting models under the same conditions.
The number of models supporting dimensionality reduction (leading-dimension preserving) differs by dimension: ( models), (), (), (), (), (), (), each with coverage over all tasks. Combination variants of quantization and dimensionality reduction also align the same coverage over all tasks for supporting models. Because these variant rows are stored side by side in the result table, the “delta vs. base” column directly shows quality degradation, and one can compare under the same conditions which models are strong at dimensions and how much int8/binary quantization degrades performance. A figure comparing quality degradation from quantization (per-model macro delta of int8/binary) and the retention rate of dimensionality reduction in native-dimension ratio, across all models, is in Appendix F.4 (Figure 8). Variant naming/correspondence and rescore details are in Appendix E.1.
Performance change from sparse-representation pruning
The evaluation targets include learned sparse models (naver/splade-v3 (SPLADE-v3; Lassance et al., 2024), prithivida/Splade_PP_en_v2, ibm-granite/granite-embedding-30m-sparse, opensearch-project/opensearch-neural-sparse-encoding-multilingual-v1). For the SPLADE family (Formal et al., 2021), we pruned with combinations of query-side and document-side max active dims, , and measured the performance change. For naver/splade-v3, the average score decreases monotonically from at to at . The document side shows almost no improvement beyond dimensions (– for ), whereas query-side reduction is more sensitive (– for at the same ). This shows the practical compression headroom: aggressive document-side pruning (reducing memory and inverted-index size) barely harms quality, while cutting the query side below has a large quality cost. The full pruning grid in base ratio and the operating envelope that keeps quality ( and ) are in Appendix F.6 (Table 10). Note that the SPLADE family is English-centric (from MS MARCO), so its average over all tasks including multilingual tasks comes out low, and absolute comparisons should be read with language held fixed. Pruning details are in Appendix E.2.
Analysis of reranking and the candidate set
Each task’s evaluation carries diagnostic records for analyzing reranker and candidate-set behavior. The default candidate set is the hybrid candidate set (top , §3.3) fusing the BM25 and dense-retrieval tops via RRF, with the safeguard that appends one positive at the tail for any query that contains none. The records include the base and reranker scores and the improvement, the candidate-set origin, query coverage (fraction of queries with at least one relevant document), relevant-document coverage (fraction of relevant documents in the top candidates), and the runtime breakdown. On the -dense-model average, query coverage was and relevant-document coverage . While the safeguard ensures every query has at least one relevant document, about of all relevant documents do not reach the top candidates.
Retrieval models re-evaluating their own candidate set.
When a dense model re-evaluates its own hybrid candidate set, the improvement is small: points on dense average ( without the safeguard metric). It is small because the hybrid candidate set already contains the dense-retrieval top, so the model rescores almost exactly the documents it ranked at the top under full-corpus retrieval. The remaining small improvement comes from the BM25-derived candidates (lexical-match documents the dense model missed under full-corpus retrieval) entering the search target, and from the safeguard always including a positive. An advantage of sharing the hybrid candidate set is that the large apparent improvement arising from the compatibility between a BM25-only candidate set and dense, which occurs when the candidate set is BM25-only, is unlikely to be mixed in.
Reranker evaluation.
The rerankers ( cross-encoders, LLM-style; §4.2) all align re-ordering results over the fixed candidate set on all tasks. Most of these rerankers—especially the encoder cross-encoders—are trained mainly for the general retrieval task of finding semantically close documents for short queries, and are most effective on tasks matching that assumption (LLM-style rerankers such as Qwen3-Reranker are an exception, as shown below). Indeed, restricting to tasks where both query and document are short (query chars and document chars), the multilingual cross-encoder BAAI/bge-reranker-v2-m3 (BGE-M3 base; Chen et al., 2024) reaches macro on short multilingual tasks, above the best dense scored directly as a reranker on the candidate set (), and the English-only cross-encoder cross-encoder/ettin-reranker-400m-v1 (Ettin; Weller et al., 2025) reaches macro on short English tasks, above the best dense (). That is, in the “short query/document retrieval” use case rerankers assume, using a reranker suited to multilingual or English respectively improves quality.
On the other hand, over all tasks including code, reasoning, instruction-following, long documents, and languages, the only reranker that exceeds the best dense (jinaai/jina-embeddings-v5-text-small, ) in reranking macro over the candidate set is the LLM-style Qwen/Qwen3-Reranker-0.6B (Zhang et al., 2025b, Zhang Y. et al.) at ; classical multilingual cross-encoders (BAAI/bge-reranker-v2-m3 , Alibaba-NLP/gte-multilingual-reranker-base (mGTE; Zhang et al., 2024) , etc.) all fall slightly below the best dense. This is because many of these rerankers are trained (often exclusively) for the short semantic-search queries above and do not generalize as broadly as the best dense to the full diversity of this benchmark.
The top reranker () also exceeds the top full-corpus dense model (), showing that reranking over the hybrid candidate set functions as a configuration that surpasses full-corpus dense retrieval under the fixed-candidate reranking protocol (with the safeguard; this is not an end-to-end production retrieval comparison, §6.4). A detailed decomposition of rerankers by type, scope, and query type (-score comparison; Table 8, Figures 3, 4) is in Appendix E.4. Note these reranker scores are computed under the premise that the safeguard includes a relevant document in the candidate set for every query, so the “degradation when the candidate set contains no positive” that real two-stage retrieval faces is isolated.
We summarize the per-scope best models in Table 2. What exceeds the best dense (jinaai/jina-embeddings-v5-text-small scored as a reranker on the same candidate set) is: over all tasks, only the LLM-style Qwen/Qwen3-Reranker-0.6B (); on the short-multilingual scope (query chars and document chars, non-English), the multilingual cross-encoder BAAI/bge-reranker-v2-m3 (; Qwen3-Reranker also exceeds it slightly at ); and on the short-English scope (same condition, English), the English-only cross-encoder cross-encoder/ettin-reranker-400m-v1 (). Thus whether a reranker beats dense depends not on “rerankers in general” but on the scope and reranker type.
| Model | Type | All 551 | Short ML | Short EN |
|---|---|---|---|---|
| Qwen3-Reranker-0.6B | LLM reranker | 68.03 | 66.48 | 66.59 |
| bge-reranker-v2-m3 | multilingual CE | 63.07 | 67.41 | 63.29 |
| ettin-reranker-400m | English-only CE | 60.82 | 58.91 | 70.23 |
| jina-v5-small | dense (reference) | 65.51 | 65.91 | 68.59 |
Per-benchmark advantage/disadvantage.
The “reranker top dense top” gap is large on multilingual, expert-domain, and reasoning benchmarks (e.g., NanoMLDR ), and even on benchmarks where the reranker top falls below, the downside is small. The per-benchmark breakdown and the decomposition of rerankers by type (cross-encoder / LLM-style), scope, and query type are organized in Appendix E.4, where we show that multilingual cross-encoders are strong on short factual queries and collapse on long queries, while LLM-style rerankers are robust to length.
Rank correlation with MTEB / MMTEB retrieval
To empirically show how well Nano-sets reproduce the ranking of the original benchmarks, we independently compared NanoMMTEB-v2, NanoMTEB-v2, and NanoBEIR-en against, respectively, MMTEB v2 retrieval, MTEB retrieval v2, and English BEIR (full) from the official mteb/results (commit 1e8ab5d, reflected up to 2026-06-08). The analysis uses only the same base rows of the §5 results on the Nano side, and excludes from the common model set any model for which the official side does not have all tasks as a single-revision single measurement, isolating pure ranking reproducibility. The aggregation assigns a rank by descending score within a task (ties get the average rank) and computes the overall ranking by averaging each model’s Borda score ( = number of models) over all tasks. The results are in Table 3, and the scatter of official vs. Nano ranks is in Figure 1.
| Metric | MMTEB | MTEB-v2 | BEIR-en |
|---|---|---|---|
| Common models | 24 | 18 | 19 |
| Tasks | 18 | 10 | 13 |
| Spearman rank correlation | 0.975 | 0.983 | 0.973 |
| Spearman 95% CI (model bootstrap) | [0.915, 0.995] | [0.912, 0.998] | [0.882, 0.997] |
| Pearson correlation (Borda score) | 0.969 | 0.981 | 0.974 |
| Mean absolute rank difference | 1.208 | 0.722 | 0.895 |
| Median rank difference | 1.000 | 1.000 | 0.500 |
| Max rank difference | 4.000 | 2.000 | 3.000 |
For all three pairs, the Spearman rank correlation exceeds , with rank differences of about on average and at most (MMTEB), (BEIR-en), and (MTEB-v2). Because the common model counts (//) are limited, we obtained confidence intervals for Spearman by bootstrap ( resamples with replacement) over the common model set (Table 3). Even at the interval lower bounds, the correlation stays at for MMTEB and MTEB-v2 and for BEIR-en, all high. The Borda-score Pearson correlation is also , so even from the perspective of aggregating per-task wins/losses, every Nano-set faithfully reproduces the official overall ranking. In particular, for both MMTEB v2 retrieval and MTEB retrieval v2, the top model is rank on both the official and Nano sides (rank difference ), a representative agreement on top-rank reproduction. Rank swaps exist, but no large movement crossing the boundaries of the top/middle/bottom groups is observed, and they are not of a scale that changes model-selection judgments.
The per-model ranking tables, per-task mean/variance differences, and the discussion of the factors behind the differences are gathered in Appendix D.
From these results, Nano-sets are not a final evaluation replacing the official full retrieval, nor do they guarantee absolute-score agreement. However, for iterative ranking judgments such as model selection, separating the top from the middle group, and pre-release regression detection, they function as a proxy that provides conclusions close to the official full evaluation at low cost, as confirmed through three independent comparisons. Fine distinctions between nearby models and conclusions that depend on a particular task still require referring to the official full tasks.
Real-data use cases
The overall ranking answers only “which model is best on average,” but practical model adoption is made under conditions such as target language, document length, latency budget, and index size. Because the benchmark measures many models tasks architectures efficiency settings under the same conditions, it directly answers such conditional adoption decisions. For example, on a pool of first-stage retrieval systems (dense , learned sparse , BM25 ), contrasting each scope’s top- system with its overall macro rank: for code RAG, instruction-following, and medical reasoning, the overall-rank- (jinaai/jina-embeddings-v5-text-small) is also the scope top-; whereas for multilingual semantic search (NanoMIRACL) the overall-rank- BAAI/bge-m3, for the two long-document series (NanoMLDR, NanoLongEmbed) the overall-rank- BM25, and for Japanese (NanoJMTEB-v2) the overall-rank- Japanese-specialized model cl-nagoya/ruri-v3-310m (Ruri; Tsukagoshi and Sasano, 2024) are each top-. Scopes where the overall score is a good guide coexist with scopes where it is a wrong guide, and which is which can only be determined by per-scope measurement. The full picture of per-scope ranks is in Appendix F.1 (Figure 5).
The observation that “the overall best model is not necessarily indicated, and the best model/architecture changes with the target scope” generalizes to three questions, each detailed as a real-data use case in Appendix F.
First, which model/architecture to choose depends on the target scope (Appendix F.1, F.2). Changing scope swaps the best model not only among dense models but also across architectures (dense / sparse / late interaction, etc.). For example, restricting to English BEIR, late interaction—not top overall—takes first place, and learned sparse enters the top quartile.
Second, different architectures can be compared on the same footing (Appendix F.3). Scoring all models as rerankers over the same fixed candidate set lets embedding models and rerankers be placed side by side. On the overall macro, only one modern general reranker exceeds the dense top, and the advantage of multilingual cross-encoders concentrates in the multilingual semantic-search scope (§5.5, Appendix E.4).
Third, the cost of efficiency settings can be read separately from quality (Appendix F.4–F.6). Dimensionality reduction and int8 quantization are predictable small costs, robustness to binary quantization depends on a model’s training characteristics, float rescoring nearly preserves cross-model comparison, and sparse pruning has a cheap document-side knob and an expensive query-side knob—each setting’s cost can be evaluated separately.
All three are material for adoption decisions that cannot be read from a single overall score, and can be extracted only when a single harness, a single task format, same-condition measurement, and a consistent aggregation basis (macro as primary in this paper; §3, §4.5) are aligned.
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org