GILTIA
GILTIA
News | Press
People
Events
Publications
Contact
English
English
Español
Low-Resource Languages
Whisper-LM con léxico maya: corrección de transcripciones ASR de maya yucateco usando Gemini y un diccionario como referencia
Un pipeline en Google Colab que corrige transcripciones ruidosas de Whisper en maya yucateco combinando reglas morfológicas, un léxico de 2,988 entradas y Gemini 2.0 Flash como modelo corrector.
Jaziel A. Carballo Tadeo
Jul 9, 2026
Whisper-LM with a Maya Lexicon: Correcting Yucatec Maya ASR Transcriptions Using Gemini and a Dictionary as Reference
A Google Colab pipeline that corrects noisy Whisper transcriptions of Yucatec Maya by combining morphological rules, a 2,988-entry lexicon, and Gemini 2.0 Flash as a corrector model.
Jaziel A. Carballo Tadeo
The YUA-ES Communicative Contexts Corpus: An Open Parallel Dataset of Everyday Yucatec Maya–Spanish Phrases
An openly licensed parallel corpus of 14,332 aligned phrase pairs in Yucatec Maya and Spanish, organized into 31 everyday communicative contexts — the first open, phrase-aligned corpus of its kind.
Alejandro Molina Villegas
,
María Elisa Chavarrea-Chim
,
Samuel Canul-Yah
,
Sary Lorena Hau-Ucán
Jun 17, 2026
PDF
Cite
DOI
View on Research Square
The YUA-ES Communicative Contexts Corpus: An Open Parallel Dataset of Everyday Yucatec Maya–Spanish Phrases
An openly licensed parallel corpus of 14,332 aligned phrase pairs in Yucatec Maya and Spanish, organized into 31 everyday communicative contexts — the first open, phrase-aligned corpus of its kind.
Alejandro Molina Villegas
,
María Elisa Chavarrea-Chim
,
Samuel Canul-Yah
,
Sary Lorena Hau-Ucán
PDF
Cite
DOI
View on Research Square
Generating a Culturally and Linguistically Adapted Word Similarity Benchmark for Yucatec Maya
Un benchmark de similitud de palabras cultural y lingüísticamente adaptado al maya yucateco, construido a partir de una Lista de Swadesh filtrada, que muestra que la elección del benchmark pesa más que el ajuste de hiperparámetros al evaluar word embeddings.
Alejandro Molina Villegas
,
Joel Suro-Villalobos
,
Jorge Reyes-Magaña
,
Silvia Fernández-Sabido
Sep 25, 2025
DOI
Ver en Inteligencia Artificial
Generating a Culturally and Linguistically Adapted Word Similarity Benchmark for Yucatec Maya
A culturally and linguistically adapted word similarity benchmark for Yucatec Maya, built from a filtered Swadesh List, showing that benchmark selection matters more than hyperparameter tuning when evaluating word embeddings.
Alejandro Molina Villegas
,
Joel Suro-Villalobos
,
Jorge Reyes-Magaña
,
Silvia Fernández-Sabido
DOI
View at Inteligencia Artificial
Mayasoundex: A Phonetically Grounded Algorithm for Information Retrieval in the Maya Language
Un algoritmo fonético tipo Soundex diseñado para la lengua maya, que genera códigos consistentes para palabras de sonido similar y permite una recuperación de información robusta frente a la variación ortográfica.
Alejandro Molina Villegas
Jun 16, 2024
PDF
Mayasoundex: A Phonetically Grounded Algorithm for Information Retrieval in the Maya Language
A Soundex-style phonetic algorithm designed for the Maya language, generating consistent codes for similar-sounding words to enable robust information retrieval despite orthographic variation.
Alejandro Molina Villegas
PDF
Cite
×