Article thumbnail

A method for measuring verb similarity for two closely related languages with application to Zulu and Xhosa

By Zola Mahlaza and C. Maria Keet

Abstract

There are limited computational resources for Nguni languages and when improving availability for one of the languages, bootstrapping from a related language’s resources may be a cost-saving approach. This requires the ability to quantify similarity between any two closely related languages so as to make informed decisions, of which it is unclear how to measure it. We devised a method for quantifying similarity by adapting four extant similar measures, and present a method of quantifying the ratio of verbs that would need phonological conditioning due to consecutive vowels. The verbs selected are those relevant for weather forecasts for Xhosa and Zulu and newly specified as computational grammar rules. The 52 Xhosa and 49 Zulu rules share 42 rules, supporting informal impressions of their similarity. The morphosyntactic similarity reached 59.5% overall on the adapted Driver-Kroeber metric, with past tense rules only at 99.5%. This similarity score is a result of the variation in terminals mainly for the prefix of the verb

Topics: Natural language generation, Language resources
Year: 2019
OAI identifier: oai:pubs.cs.uct.ac.za:1360

Suggested articles


To submit an update or takedown request for this paper, please submit an Update/Correction/Removal Request.