Russian Information Retrieval Benchmark

rusBEIR-lite Leaderboard (23/29)

Compare dense retrievers, sparse baselines, and reranker pipelines on official rusBEIR datasets. The default ranking is the macro-average of NDCG@10.

Model Rankings

Filter by task family, model name, or verification status. Scores are stored as fractions; leaderboard rankings use the selected average metric.

Datasets (23/29)

rusBEIR-lite tasks used for the default macro-average ranking.

About rusBEIR-lite

rusBEIR-lite — облегчённая версия полного бенчмарка rusBEIR. Для сокращения времени тестирования исключены задачи, выполнение которых занимало более 60 минут для BGE-Reranker-v2-m3.

Пропуск 6 наиболее затратных по времени задач из 29 (rus-mmarco, rus-miracl, rus-cqadupstack, ria-news, wikifacts-articles, wikifacts-para) сократил время тестирования примерно в 3 раза — до 10–12 часов на модель.

About rusBEIR

rusBEIR is a Russian BEIR-style benchmark for zero-shot information retrieval. The leaderboard is backed by a plain JSONL file, so every row can be reviewed or mirrored to a Hugging Face Dataset.

Verified rows should point to reproducible logs or a commit with generated retrieval results.

Project repository: kaengreg/rusBEIR

Citation

@inproceedings{kovalev2025building,
  title={Building Russian Benchmark for Evaluation of Information Retrieval Models},
  author={Kovalev, Grigory and Tikhomirov, Mikhail and Kozhevnikov, Evgeny and Kornilov, Max and Loukachevitch, Natalia},
  booktitle={Proceedings of the International Conference "Dialogue},
  volume={2025},
  year={2025}
}