Compare dense retrievers, sparse baselines, and reranker pipelines on official rusBEIR datasets. The default ranking is the macro-average of NDCG@10.
Filter by task family, model name, or verification status. Scores are stored as fractions; leaderboard rankings use the selected average metric.
rusBEIR-lite tasks used for the default macro-average ranking.
rusBEIR-lite — облегчённая версия полного бенчмарка rusBEIR. Для сокращения времени тестирования исключены задачи, выполнение которых занимало более 60 минут для BGE-Reranker-v2-m3.
Пропуск 6 наиболее затратных по времени задач из 29 (rus-mmarco, rus-miracl, rus-cqadupstack, ria-news, wikifacts-articles, wikifacts-para) сократил время тестирования примерно в 3 раза — до 10–12 часов на модель.
rusBEIR is a Russian BEIR-style benchmark for zero-shot information retrieval. The leaderboard is backed by a plain JSONL file, so every row can be reviewed or mirrored to a Hugging Face Dataset.
Verified rows should point to reproducible logs or a commit with generated retrieval results.
Project repository: kaengreg/rusBEIR
@inproceedings{kovalev2025building,
title={Building Russian Benchmark for Evaluation of Information Retrieval Models},
author={Kovalev, Grigory and Tikhomirov, Mikhail and Kozhevnikov, Evgeny and Kornilov, Max and Loukachevitch, Natalia},
booktitle={Proceedings of the International Conference "Dialogue},
volume={2025},
year={2025}
}