The YUA-ES Communicative Contexts Corpus: An Open Parallel Dataset of Everyday Yucatec Maya–Spanish Phrases

Abstract

We present the YUA-ES Communicative Contexts Corpus (YUA-ES-CCC), an openly licensed parallel corpus of 14,332 aligned phrase pairs in Yucatec Maya (Glottolog yuca1254) and Spanish, organized into 31 everyday communicative contexts. To the best of our knowledge, YUA-ES-CCC is the first openly licensed, everyday-register, phrase-aligned Yucatec Maya–Spanish corpus. We describe the corpus and its schema, the collaborative creation process under a formal institutional agreement, a reproducible consolidation and validation pipeline, and descriptive statistics, and we outline its reuse potential for machine translation, educational-material generation, and the linguistic study of everyday Yucatec Maya. Released under CC-BY-4.0 with a persistent DOI, a machine-readable citation file, and a datasheet, YUA-ES-CCC shows that well-curated, culturally pertinent open data—built with and credited to the speaker community—is itself a meaningful contribution to language technology for Yucatec Maya and other low-resource languages.

Publication
Research Square (preprint)
Alejandro Molina Villegas
Alejandro Molina Villegas
Languages, Territory and AI Group Lead

Cátedras Conacyt researcher at the Centro de Investigación en Ciencias de Información Geoespacial (CentroGeo), Yucatán. His research lines include natural language processing, data mining, geoparsing, machine learning, and technologies for the Maya language.