Generating a Culturally and Linguistically Adapted Word Similarity Benchmark for Yucatec Maya

Abstract

We propose a novel methodology for constructing reliable word embeddings for Yucatec Maya by adapting the Swadesh List for semantic similarity evaluation. The linguistic resource was translated into Yucatec Maya and refined through cultural and linguistic filtering, and similarity scores between embeddings were correlated against the resulting benchmark. Our findings indicate that the selection of the evaluation benchmark significantly outweighs hyperparameter optimization in determining performance outcomes, highlighting the importance of culturally grounded evaluation resources for low-resource NLP.

Publication
Inteligencia Artificial, 28(76), 283–300
Alejandro Molina Villegas
Alejandro Molina Villegas
Languages, Territory and AI Group Lead

Cátedras Conacyt researcher at the Centro de Investigación en Ciencias de Información Geoespacial (CentroGeo), Yucatán. His research lines include natural language processing, data mining, geoparsing, machine learning, and technologies for the Maya language.