The YUA-ES Communicative Contexts Corpus: An Open Parallel Dataset of Everyday Yucatec Maya–Spanish Phrases

Resumen

We present the YUA-ES Communicative Contexts Corpus (YUA-ES-CCC), an openly licensed parallel corpus of 14,332 aligned phrase pairs in Yucatec Maya (Glottolog yuca1254) and Spanish, organized into 31 everyday communicative contexts. To the best of our knowledge, YUA-ES-CCC is the first openly licensed, everyday-register, phrase-aligned Yucatec Maya–Spanish corpus. We describe the corpus and its schema, the collaborative creation process under a formal institutional agreement, a reproducible consolidation and validation pipeline, and descriptive statistics, and we outline its reuse potential for machine translation, educational-material generation, and the linguistic study of everyday Yucatec Maya. Released under CC-BY-4.0 with a persistent DOI, a machine-readable citation file, and a datasheet, YUA-ES-CCC shows that well-curated, culturally pertinent open data—built with and credited to the speaker community—is itself a meaningful contribution to language technology for Yucatec Maya and other low-resource languages.

Publicación
Research Square (preprint)
Alejandro Molina Villegas
Alejandro Molina Villegas
Líder del Grupo de Investigación en Lenguas, Territorio e Inteligencia Artificial

Investigador por México (Cátedras Conacyt) en el Centro de Investigación en Ciencias de Información Geoespacial (CentroGeo), subsede Yucatán. Sus líneas de investigación incluyen procesamiento de lenguaje natural, minería de datos, geoparsing, aprendizaje de máquina y tecnologías para la lengua maya.