This study introduces a benchmark for evaluating information retrieval in the Slovak language, leveraging a question-answering dataset for fine-tuning and assessment of sentence transformers. The dataset, named Retrieval SkQuAD, is integrated into two widely recognized evaluation frameworks: BEIR (Benchmarking Information Retrieval) and MTEB (Massive Text Embedding Benchmark). Retrieval SkQuAD (Slovak Question Answering Dataset) comprises 19,000 manually annotated answers to 1,134 questions, with each answer assigned a relevance score, and includes information on the usefulness of documents in generating responses. Unlike question answering datasets, this resource provides a nuanced assessment of partial relevance across multiple documents. We fine-tuned several sentence transformers and BERT-based models specifically for retrieving documents containing correct answers within the Slovak Wikipedia. Our fine-tuning process incorporates adversarial questions as hard negatives, leading to significant improvements in retrieval accuracy. Experimental results demonstrate that our approach advances state-of-the-art performance for Slovak information retrieval.
Hládek et al. (2026) studied this question.