Slovene Natural Language Inference Dataset SI-NLI

Klemen, Matej; Robnik-Šikonja, Marko; Čibej, Jaka; Žagar, Aleš

Slovene Natural Language Inference Dataset SI-NLI

Authors: Matej Klemen
Marko Robnik-Šikonja
Jaka Čibej
Aleš Žagar
Publication date: 13 November 2022
Publisher: University of Ljubljana

Abstract

SI-NLI (Slovene Natural Language Inference Dataset) contains 5,937 human-created Slovene sentence pairs (premise and hypothesis) that are manually labeled with the labels "entailment", "contradiction", and "neutral". We created the dataset using sentences that appear in the Slovenian reference corpus ccKres (http://hdl.handle.net/11356/1034). Annotators were tasked to modify the hypothesis in a candidate pair in a way that reflects one of the labels. The dataset is balanced since the annotators created three modifications (entailment, contradiction, neutral) for each candidate sentence pair. The dataset is split into train, validation, and test sets, with sizes of 4,392, 547, and 998. We used Slovenian pre-trained language models to create splits, thereby ensuring that difficult and easy instances are evenly distributed in all three subsets. The dataset is released in a tabular TSV format. The README.txt file contains a description of the attributes. Only the hypothesis and premise are given in the test set (i.e. no annotations) since SI-NLI is integrated into the Slovene evaluation framework SloBENCH (https://slobench.cjvt.si/). If you use the dataset to train your models, please consider submitting the test set predictions to SloBENCH to get the evaluation score and see how it compares to others

Similar works

Full text

Available Versions

Common Language Resources and Technology Infrastructure - Slovenia

oai:www.clarin.si:11356/1707

Last time updated on 19/11/2022