Slovene and Croatian word embeddings in terms of gender occupational analogies

Anka Supej; Marko Robnik-Šikonja; Matej Ulčar; Senja Pollak

Slovene and Croatian word embeddings in terms of gender occupational analogies

Authors: Anka Supej
Marko Robnik-Šikonja
Matej Ulčar
Senja Pollak
Publication date: 1 July 2021
Publisher: 'University of Ljubljana'
Doi

Abstract

In recent years, the use of deep neural networks and dense vector embeddings for text representation have led to excellent results in the field of computational understanding of natural language. It has also been shown that word embeddings often capture gender, racial and other types of bias. The article focuses on evaluating Slovene and Croatian word embeddings in terms of gender bias using word analogy calculations. We compiled a list of masculine and feminine nouns for occupations in Slovene and evaluated the gender bias of fastText, word2vec and ELMo embeddings with different configurations and different approaches to analogy calculations. The lowest occupational gender bias was observed with the fastText embeddings. Similarly, we compared different fastText embeddings on Croatian occupational analogies

Similar works

Full text

Open in the Core reader

Download PDF

Available Versions

Directory of Open Access Journals

oai:doaj.org/article:9d4e22fbe...

Last time updated on 26/01/2023