Generating a Culturally and Linguistically Adapted Word Similarity Benchmark for Yucatec Maya

Resumen

We propose a novel methodology for constructing reliable word embeddings for Yucatec Maya by adapting the Swadesh List for semantic similarity evaluation. The linguistic resource was translated into Yucatec Maya and refined through cultural and linguistic filtering, and similarity scores between embeddings were correlated against the resulting benchmark. Our findings indicate that the selection of the evaluation benchmark significantly outweighs hyperparameter optimization in determining performance outcomes, highlighting the importance of culturally grounded evaluation resources for low-resource NLP.

Publicación
Inteligencia Artificial, 28(76), 283–300
Alejandro Molina Villegas
Alejandro Molina Villegas
Líder del Grupo de Investigación en Lenguas, Territorio e Inteligencia Artificial

Investigador por México (Cátedras Conacyt) en el Centro de Investigación en Ciencias de Información Geoespacial (CentroGeo), subsede Yucatán. Sus líneas de investigación incluyen procesamiento de lenguaje natural, minería de datos, geoparsing, aprendizaje de máquina y tecnologías para la lengua maya.