<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Languages, Territory and AI Research Group</title><link>https://giltia.github.io/</link><atom:link href="https://giltia.github.io/index.xml" rel="self" type="application/rss+xml"/><description>Languages, Territory and AI Research Group</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Mon, 01 Jan 2024 00:00:00 +0000</lastBuildDate><image><url>https://giltia.github.io/media/icon_hu_4650b61b0e05ba9e.png</url><title>Languages, Territory and AI Research Group</title><link>https://giltia.github.io/</link></image><item><title>Diario de Yucatán: Traductor en maya</title><link>https://giltia.github.io/es/post/traductor-en-maya-diario-yucatan/</link><pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/post/traductor-en-maya-diario-yucatan/</guid><description>&lt;p&gt;&lt;strong&gt;Perfiles de nuestra gente&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="traductor-en-maya"&gt;Traductor en maya&lt;/h2&gt;
&lt;h3 id="buscan-integrar-la-lengua-regional-a-la-tecnología"&gt;Buscan integrar la lengua regional a la tecnología&lt;/h3&gt;
&lt;p&gt;La investigación desarrollada en Yucatán para acercar la lengua maya a las nuevas tecnologías llegó a uno de los escenarios académicos más importantes del mundo en materia de inteligencia artificial.&lt;/p&gt;
&lt;p&gt;El doctorante del CentroGeo, Jaziel Aarón Carballo-Tadeo, representó a México en la 16.ª edición de la Escuela de Verano de Aprendizaje Automático de Lisboa (LxMLS 2026), donde presentó un proyecto que busca enseñar a las computadoras a comprender y transcribir el maya yucateco.&lt;/p&gt;
&lt;p&gt;El objetivo de la investigación de Carballo-Tadeo fue reducir la brecha tecnológica que enfrentan los hablantes de lenguas originarias, ya que actualmente los sistemas de reconocimiento de voz funcionan de manera eficiente en idiomas como el español o el inglés, pero prácticamente no reconocen el maya yucateco.&lt;/p&gt;
&lt;p&gt;Él mencionó que la investigación propone un método para el mejoramiento de la precisión con la que las computadoras transcriben la voz de los hablantes de esta lengua, mediante el uso de diccionarios especializados y las reglas oficiales de escritura del maya.&lt;/p&gt;
&lt;p&gt;Según destacó, haber sido admitido a la LxMLS 2026 representa un importante reconocimiento, porque se trata de una convocatoria internacional altamente competitiva que reúne cada año a estudiantes e investigadores de distintos países, así como a especialistas de instituciones de prestigio como la Universidad de Washington, Carnegie Mellon y el Allen Institute for AI.&lt;/p&gt;
&lt;p&gt;Además de ser de los seleccionados para participar en este encuentro académico, su investigación fue una de las 24 seleccionadas para presentarse en formato de póster, lo que le permitió compartir ante la comunidad científica internacional un proyecto desarrollado desde CentroGeo sede Mérida y enfocado en atender las necesidades tecnológicas de la lengua maya.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Idioma invisible&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;El doctorante del CentroGeo, Jaziel Carballo-Tadeo, indicó que el idioma maya sigue siendo invisible para la mayoría de los asistentes virtuales, teléfonos inteligentes y plataformas de reconocimiento de voz.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Presencia limitada&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Además, subrayó que la presencia de investigadores latinoamericanos en este tipo de encuentros internacionales, como la Escuela de Verano de Aprendizaje Automático de Lisboa, aún es limitada.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Paso importante&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;La participación de una investigación centrada en enseñar a las computadoras a comprender y transcribir el maya yucateco representa un paso importante para posicionar al idioma dentro del desarrollo de las tecnologías del lenguaje y fortalecer futuras colaboraciones científicas internacionales.&lt;/p&gt;
&lt;p&gt;El investigador dijo que este proyecto fue realizado en colaboración con el doctor Alejandro Molina-Villegas, responsable del desarrollo de tecnologías lingüísticas en CentroGeo Mérida, con el propósito de impulsar herramientas que permitan la inclusión digital de las lenguas indígenas.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Nota de Ilse Noh Canché, publicada en Diario de Yucatán el 17 de agosto de 2026.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Diario de Yucatán: Traductor en maya ("Translator in Maya")</title><link>https://giltia.github.io/post/traductor-en-maya-diario-yucatan/</link><pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/post/traductor-en-maya-diario-yucatan/</guid><description>&lt;p&gt;&lt;strong&gt;Profiles of our people&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="translator-in-maya"&gt;Translator in Maya&lt;/h2&gt;
&lt;h3 id="seeking-to-bring-the-regional-language-into-technology"&gt;Seeking to bring the regional language into technology&lt;/h3&gt;
&lt;p&gt;Research developed in Yucatán to bring the Maya language closer to new technologies reached one of the world&amp;rsquo;s most important academic stages in artificial intelligence.&lt;/p&gt;
&lt;p&gt;CentroGeo PhD candidate Jaziel Aarón Carballo-Tadeo represented Mexico at the 16th Lisbon Machine Learning School (LxMLS 2026), where he presented a project aimed at teaching computers to understand and transcribe Yucatec Maya.&lt;/p&gt;
&lt;p&gt;The goal of Carballo-Tadeo&amp;rsquo;s research was to reduce the technological gap faced by speakers of Indigenous languages, since current speech-recognition systems work efficiently for languages such as Spanish or English but barely recognize Yucatec Maya.&lt;/p&gt;
&lt;p&gt;He noted that the research proposes a method to improve the accuracy with which computers transcribe the speech of Yucatec Maya speakers, through the use of specialized dictionaries and the official rules of Maya orthography.&lt;/p&gt;
&lt;p&gt;He highlighted that being admitted to LxMLS 2026 is a significant recognition, since it is a highly competitive international call that brings together students and researchers from different countries each year, along with specialists from prestigious institutions such as the University of Washington, Carnegie Mellon, and the Allen Institute for AI.&lt;/p&gt;
&lt;p&gt;In addition to being selected to take part in the academic gathering, his research was one of 24 selected for poster presentation, allowing him to share with the international scientific community a project developed at CentroGeo&amp;rsquo;s Mérida campus and focused on addressing the technological needs of the Maya language.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An invisible language&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;CentroGeo PhD candidate Jaziel Carballo-Tadeo said the Maya language remains invisible to most virtual assistants, smartphones, and voice-recognition platforms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Limited presence&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;He also stressed that the presence of Latin American researchers at international gatherings such as the Lisbon Machine Learning School is still limited.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An important step&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Presenting research focused on teaching computers to understand and transcribe Yucatec Maya marks an important step toward positioning the language within language-technology development and strengthening future international scientific collaborations.&lt;/p&gt;
&lt;p&gt;The researcher said the project was carried out in collaboration with Dr. Alejandro Molina-Villegas, who leads the development of language technologies at CentroGeo Mérida, with the goal of advancing tools that enable the digital inclusion of Indigenous languages.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Article by Ilse Noh Canché, published in Diario de Yucatán on August 17, 2026. Translated from the original Spanish for this archive.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>TecNM Campus Mérida: orgullo por la investigación en IA y lengua maya de un egresado</title><link>https://giltia.github.io/es/post/tecnm-campus-merida-ia-lengua-maya/</link><pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/post/tecnm-campus-merida-ia-lengua-maya/</guid><description>&lt;p&gt;El Tecnológico Nacional de México (TecNM), Campus Mérida, publicó en su página de Facebook un mensaje de reconocimiento a Jaziel Aarón Carballo Tadeo, egresado de Ingeniería en Sistemas Computacionales, por impulsar el desarrollo de tecnologías de inteligencia artificial capaces de reconocer y procesar la lengua maya yucateca, trabajo presentado en la 16.ª Escuela de Verano de Aprendizaje Automático de Lisboa (LxMLS 2026), en Portugal. La institución subrayó que el proyecto combina innovación tecnológica con la preservación lingüística y el patrimonio cultural.&lt;/p&gt;
&lt;div style="text-align:center;"&gt;
&lt;iframe src="https://www.facebook.com/plugins/post.php?href=https%3A%2F%2Fwww.facebook.com%2FTecNMCampusMerida%2Fposts%2Fpfbid0VnE9rjJUKpb36Dbmxm7DY9zNgCgW8RAj5GbXB4Hr6veEyEKtgEcELMUshqSqQp4Sl&amp;show_text=true&amp;width=500" width="500" height="709" style="border:none;overflow:hidden;max-width:100%;" scrolling="no" frameborder="0" allowfullscreen="true" allow="autoplay; clipboard-write; encrypted-media; picture-in-picture; web-share"&gt;&lt;/iframe&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Publicación de TecNM Campus Mérida en Facebook, 16 de agosto de 2026.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>TecNM Campus Mérida: pride in an alumnus's AI and Maya-language research</title><link>https://giltia.github.io/post/tecnm-campus-merida-ia-lengua-maya/</link><pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/post/tecnm-campus-merida-ia-lengua-maya/</guid><description>&lt;p&gt;Tecnológico Nacional de México (TecNM), Campus Mérida, posted a message on its Facebook page recognizing its alumnus Jaziel Aarón Carballo Tadeo, a graduate of Computer Systems Engineering, for advancing the development of artificial intelligence technologies capable of recognizing and processing Yucatec Maya — work presented at the 16th Lisbon Machine Learning School (LxMLS 2026) in Portugal. The institution noted that the project combines technological innovation with linguistic preservation and cultural heritage.&lt;/p&gt;
&lt;div style="text-align:center;"&gt;
&lt;iframe src="https://www.facebook.com/plugins/post.php?href=https%3A%2F%2Fwww.facebook.com%2FTecNMCampusMerida%2Fposts%2Fpfbid0VnE9rjJUKpb36Dbmxm7DY9zNgCgW8RAj5GbXB4Hr6veEyEKtgEcELMUshqSqQp4Sl&amp;show_text=true&amp;width=500" width="500" height="709" style="border:none;overflow:hidden;max-width:100%;" scrolling="no" frameborder="0" allowfullscreen="true" allow="autoplay; clipboard-write; encrypted-media; picture-in-picture; web-share"&gt;&lt;/iframe&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Post by TecNM Campus Mérida on Facebook, August 16, 2026. Translated from the original Spanish for this archive.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>TeleYucatán: Entrevista sobre LxMLS 2026 y el proyecto de lengua maya</title><link>https://giltia.github.io/es/post/teleyucatan-entrevista-lxmls-2026/</link><pubDate>Fri, 14 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/post/teleyucatan-entrevista-lxmls-2026/</guid><description>&lt;p&gt;Entrevista en vivo con Jaziel Carballo Tadeo sobre su participación en la 16.ª Escuela de Verano de Aprendizaje Automático de Lisboa (LxMLS 2026), en la que habló de cómo llegó a Lisboa, Portugal, y en qué consiste su proyecto de investigación para que las tecnologías de inteligencia artificial reconozcan y procesen la lengua maya.&lt;/p&gt;
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;"&gt;
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share; fullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/GQsOUYaQ0u8?autoplay=0&amp;amp;controls=1&amp;amp;end=0&amp;amp;loop=0&amp;amp;mute=0&amp;amp;start=2123" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"&gt;&lt;/iframe&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Entrevista transmitida en vivo por TeleYucatán, agosto de 2026.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>TeleYucatán: Interview on LxMLS 2026 and the Maya-language project</title><link>https://giltia.github.io/post/teleyucatan-entrevista-lxmls-2026/</link><pubDate>Fri, 14 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/post/teleyucatan-entrevista-lxmls-2026/</guid><description>&lt;p&gt;Live interview with Jaziel Carballo Tadeo about his participation in the 16th Lisbon Machine Learning School (LxMLS 2026), in which he talked about how he got to Lisbon, Portugal, and what his research project on artificial intelligence technologies for recognizing and processing the Maya language is about.&lt;/p&gt;
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;"&gt;
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share; fullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/GQsOUYaQ0u8?autoplay=0&amp;amp;controls=1&amp;amp;end=0&amp;amp;loop=0&amp;amp;mute=0&amp;amp;start=2123" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"&gt;&lt;/iframe&gt;
&lt;/div&gt;
&lt;p&gt;&lt;em&gt;Interview broadcast live by TeleYucatán, August 2026.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>TV Azteca: El investigador yucateco Jaziel Carballo trabaja para que la IA reconozca, transcriba y traduzca la lengua maya</title><link>https://giltia.github.io/es/post/tvazteca-hechos-meridiano-lengua-maya/</link><pubDate>Fri, 14 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/post/tvazteca-hechos-meridiano-lengua-maya/</guid><description>&lt;p&gt;&lt;strong&gt;#HechosMeridianoYucatán&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;El investigador yucateco Jaziel Carballo trabaja en un proyecto para que la Inteligencia Artificial reconozca, transcriba y traduzca con precisión la lengua maya, buscando cerrar la brecha digital y preservar esta lengua ancestral.&lt;/p&gt;
&lt;p&gt;Entrevista en video con Jaziel Carballo, identificado en el segmento como &amp;ldquo;Doctorante Ciencias en Información Geoespacial&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Segmento de TV Azteca Yucatán, transmitido el 14 de agosto de 2026 a las 22:00 h.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>TV Azteca: Yucatecan researcher Jaziel Carballo works to have AI recognize, transcribe, and translate the Maya language</title><link>https://giltia.github.io/post/tvazteca-hechos-meridiano-lengua-maya/</link><pubDate>Fri, 14 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/post/tvazteca-hechos-meridiano-lengua-maya/</guid><description>&lt;p&gt;&lt;strong&gt;#HechosMeridianoYucatán&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Yucatecan researcher Jaziel Carballo is working on a project for Artificial Intelligence to accurately recognize, transcribe, and translate the Maya language, seeking to close the digital gap and preserve this ancestral language.&lt;/p&gt;
&lt;p&gt;Video interview with Jaziel Carballo, identified in the segment as &amp;ldquo;PhD Candidate in Geospatial Information Sciences.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;TV Azteca Yucatán segment, aired August 14, 2026 at 10:00 p.m. Translated from the original Spanish for this archive.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>PuntoMedio: Buscan que la inteligencia artificial reconozca y procese la lengua maya</title><link>https://giltia.github.io/es/post/puntomedio-ia-reconozca-lengua-maya/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/post/puntomedio-ia-reconozca-lengua-maya/</guid><description>&lt;p&gt;&lt;strong&gt;Propuesta&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="buscan-que-la-inteligencia-artificial-reconozca-y-procese-la-lengua-maya"&gt;Buscan que la inteligencia artificial reconozca y procese la lengua maya&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;Texto y fotos: Andrea Segura&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;El maya yucateco llegó a un escenario internacional de investigación en inteligencia artificial (IA) de la mano de Jaziel Aarón Carballo Tadeo, quien presentó en Lisboa, Portugal, una propuesta para avanzar en el desarrollo de tecnologías capaces de procesar y reconocer esta lengua originaria.&lt;/p&gt;
&lt;p&gt;Carballo Tadeo, estudiante del Doctorado en Ciencias de Información Geoespacial del CentroGeo, participó en la 16ª Escuela de Verano de Aprendizaje Automático de Lisboa (LxMLS 2026), organizada por el Instituto Superior Técnico de la Universidad de Lisboa en colaboración con ELLIS, red europea dedicada a la investigación en IA.&lt;/p&gt;
&lt;p&gt;De acuerdo con la información proporcionada, alrededor de 500 personas de más de 20 países solicitaron ingresar al programa, de las cuales únicamente 135 fueron seleccionadas. El investigador yucateco fue el único mexicano admitido. Además de participar en el programa académico, su investigación se centró particularmente en tecnologías de voz y reconocimiento del lenguaje. Para ello, la propuesta utiliza diccionarios y las reglas oficiales de escritura de la lengua como parte de los recursos para desarrollar modelos que permitan a las tecnologías procesarla.&lt;/p&gt;
&lt;p&gt;El proyecto busca contribuir a reducir la brecha tecnológica que enfrentan las lenguas originarias frente a idiomas con mayor representación en sistemas de inteligencia artificial, con potenciales aplicaciones en herramientas de reconocimiento de voz, asistentes digitales y otras tecnologías de procesamiento del lenguaje.&lt;/p&gt;
&lt;p&gt;Carballo Tadeo, egresado del Instituto Tecnológico de Mérida, llevó así una investigación centrada en el patrimonio lingüístico de la península de Yucatán a uno de los espacios internacionales de formación científica en inteligencia artificial.&lt;/p&gt;
&lt;p&gt;La participación internacional del investigador contó con el respaldo de la Secretaría de Ciencias, Humanidades, Tecnología e Innovación de Yucatán (Secihti), encabezada por Mirna Manzanilla Romero, institución que apoyó las gestiones necesarias para facilitar su traslado y participación en el encuentro realizado en Portugal.&lt;/p&gt;
&lt;p&gt;La propuesta plantea que el desarrollo de nuevas tecnologías no solo considere los idiomas de mayor presencia internacional, sino que también incorpore las lenguas originarias, con el objetivo de ampliar su presencia en el entorno digital y contribuir a su preservación y uso en nuevas herramientas tecnológicas.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Datos a destacar:&lt;/strong&gt; Jaziel Aarón Carballo Tadeo presentó en Portugal una propuesta para avanzar en el desarrollo de tecnologías capaces de procesar y reconocer la lengua maya.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Nota de Andrea Segura, publicada en PuntoMedio el 11 de agosto de 2026.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>PuntoMedio: Seeking to have artificial intelligence recognize and process the Maya language</title><link>https://giltia.github.io/post/puntomedio-ia-reconozca-lengua-maya/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/post/puntomedio-ia-reconozca-lengua-maya/</guid><description>&lt;p&gt;&lt;strong&gt;Proposal&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="seeking-to-have-artificial-intelligence-recognize-and-process-the-maya-language"&gt;Seeking to have artificial intelligence recognize and process the Maya language&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;Text and photos: Andrea Segura&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Yucatec Maya reached an international artificial intelligence (AI) research stage through Jaziel Aarón Carballo Tadeo, who presented a proposal in Lisbon, Portugal, to advance the development of technologies capable of processing and recognizing this Indigenous language.&lt;/p&gt;
&lt;p&gt;Carballo Tadeo, a PhD candidate in Geospatial Information Sciences at CentroGeo, took part in the 16th Lisbon Machine Learning School (LxMLS 2026), organized by the Instituto Superior Técnico of the University of Lisbon in collaboration with ELLIS, the European network for AI research.&lt;/p&gt;
&lt;p&gt;According to the information provided, around 500 people from more than 20 countries applied to the program, of whom only 135 were selected. The Yucatecan researcher was the only Mexican admitted. In addition to taking part in the academic program, his research focused particularly on speech technology and language recognition. To do so, the proposal uses dictionaries and the official orthographic rules of the language as part of the resources for developing models that allow technologies to process it.&lt;/p&gt;
&lt;p&gt;The project seeks to help close the technological gap faced by Indigenous languages compared with languages that have greater representation in artificial intelligence systems, with potential applications in speech-recognition tools, digital assistants, and other language-processing technologies.&lt;/p&gt;
&lt;p&gt;Carballo Tadeo, a graduate of the Instituto Tecnológico de Mérida, thus brought research centered on the linguistic heritage of the Yucatán Peninsula to one of the international spaces for scientific training in artificial intelligence.&lt;/p&gt;
&lt;p&gt;The researcher&amp;rsquo;s international participation was supported by Yucatán&amp;rsquo;s Secretariat of Science, Humanities, Technology and Innovation (Secihti), led by Mirna Manzanilla Romero, which backed the arrangements needed to facilitate his travel and participation in the gathering held in Portugal.&lt;/p&gt;
&lt;p&gt;The proposal argues that the development of new technologies should not only consider languages with greater international presence, but should also incorporate Indigenous languages, with the aim of expanding their presence in the digital environment and contributing to their preservation and use in new technological tools.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Highlights:&lt;/strong&gt; Jaziel Aarón Carballo Tadeo presented a proposal in Portugal to advance the development of technologies capable of processing and recognizing the Maya language.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Article by Andrea Segura, published in PuntoMedio on August 11, 2026. Translated from the original Spanish for this archive.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Radio Fórmula: ¡El maya yucateco llega a la inteligencia artificial!</title><link>https://giltia.github.io/es/post/radio-formula-ia-lengua-maya/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/post/radio-formula-ia-lengua-maya/</guid><description>&lt;p&gt;&lt;strong&gt;Radio Fórmula Yucatán&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;🇲🇽🤖 ¡El maya yucateco llega a la inteligencia artificial!&lt;/p&gt;
&lt;p&gt;El investigador yucateco Jaziel Aarón Carballo Tadeo llevó hasta Lisboa, Portugal, una investigación que busca que las computadoras puedan reconocer y procesar la lengua maya.&lt;/p&gt;
&lt;p&gt;📡 Su trabajo fue seleccionado entre solo 24 investigaciones para ser presentado en la 16.ª Escuela de Verano de Aprendizaje Automático de Lisboa (LxMLS 2026), uno de los encuentros de IA más destacados de Europa.&lt;/p&gt;
&lt;p&gt;🌎 Con más de 774 mil hablantes en la Península de Yucatán, el proyecto busca que la lengua maya también tenga presencia en el futuro tecnológico y no quede fuera de herramientas como asistentes virtuales y sistemas de reconocimiento de voz.&lt;/p&gt;
&lt;p&gt;🇲🇽 Desde Yucatán, la investigación busca abrir camino para que una de las lenguas originarias más vivas de México también tenga un lugar en la era digital.&lt;/p&gt;
&lt;p&gt;👏 ¡Orgullo yucateco que llega hasta Lisboa! @jazielcarballo&lt;/p&gt;
&lt;p&gt;#OrgulloYucateco #FormulaNoticias #lxmls2026&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Publicación de Radio Fórmula Yucatán en Facebook, agosto de 2026.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Radio Fórmula: Yucatec Maya reaches artificial intelligence!</title><link>https://giltia.github.io/post/radio-formula-ia-lengua-maya/</link><pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/post/radio-formula-ia-lengua-maya/</guid><description>&lt;p&gt;&lt;strong&gt;Radio Fórmula Yucatán&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;🇲🇽🤖 Yucatec Maya reaches artificial intelligence!&lt;/p&gt;
&lt;p&gt;Yucatecan researcher Jaziel Aarón Carballo Tadeo brought research to Lisbon, Portugal, aimed at having computers recognize and process the Maya language.&lt;/p&gt;
&lt;p&gt;📡 His work was selected among only 24 research projects to be presented at the 16th Lisbon Machine Learning School (LxMLS 2026), one of Europe&amp;rsquo;s leading AI gatherings.&lt;/p&gt;
&lt;p&gt;🌎 With more than 774,000 speakers on the Yucatán Peninsula, the project seeks to give the Maya language a place in the technological future so it is not left out of tools such as virtual assistants and speech-recognition systems.&lt;/p&gt;
&lt;p&gt;🇲🇽 From Yucatán, the research seeks to pave the way for one of Mexico&amp;rsquo;s most widely spoken Indigenous languages to also have a place in the digital era.&lt;/p&gt;
&lt;p&gt;👏 Yucatecan pride that reaches all the way to Lisbon! @jazielcarballo&lt;/p&gt;
&lt;p&gt;#OrgulloYucateco #FormulaNoticias #lxmls2026&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Post by Radio Fórmula Yucatán on Facebook, August 2026. Translated from the original Spanish for this archive.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Telesur: Investigador yucateco lleva el maya yucateco a Lisboa</title><link>https://giltia.github.io/es/post/telesur-ia-entendera-lengua-maya/</link><pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/post/telesur-ia-entendera-lengua-maya/</guid><description>&lt;p&gt;&lt;strong&gt;IA entenderá — Lengua maya gracias a investigador&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Investigador yucateco lleva el maya yucateco a Lisboa: propone método para que la inteligencia artificial entienda la lengua.&lt;/p&gt;
&lt;p&gt;Más de 774 mil hablantes de maya yucateco en la Península enfrentan una paradoja cotidiana: mientras cualquier teléfono reconoce el inglés, el español o el francés con facilidad, la lengua originaria más viva de la región es prácticamente invisible para la tecnología. Ningún asistente virtual la entiende. Ninguna aplicación la reconoce. Una investigación presentada en Lisboa, Portugal, propone un método para comenzar a cerrar esa brecha.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Publicación de Telesur (@telesuryuc) en redes sociales, agosto de 2026.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Telesur: Yucatecan researcher brings Yucatec Maya to Lisbon</title><link>https://giltia.github.io/post/telesur-ia-entendera-lengua-maya/</link><pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/post/telesur-ia-entendera-lengua-maya/</guid><description>&lt;p&gt;&lt;strong&gt;AI Will Understand — Maya Language Thanks to Researcher&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Yucatecan researcher brings Yucatec Maya to Lisbon: proposes a method for artificial intelligence to understand the language.&lt;/p&gt;
&lt;p&gt;More than 774,000 Yucatec Maya speakers on the Peninsula face a daily paradox: while any phone easily recognizes English, Spanish, or French, the region&amp;rsquo;s most widely spoken Indigenous language is practically invisible to technology. No virtual assistant understands it. No app recognizes it. Research presented in Lisbon, Portugal, proposes a method to start closing that gap.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Post by Telesur (@telesuryuc) on social media, August 2026. Translated from the original Spanish for this archive.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>LxMLS 2026: 16.ª Escuela de Aprendizaje Automático de Lisboa</title><link>https://giltia.github.io/es/event/lxmls-2026/</link><pubDate>Mon, 20 Jul 2026 09:00:00 +0000</pubDate><guid>https://giltia.github.io/es/event/lxmls-2026/</guid><description>&lt;p&gt;La &lt;a href="https://lxmls.github.io/2026/" target="_blank" rel="noopener"&gt;Lisbon Machine Learning School (LxMLS)&lt;/a&gt; celebra en 2026 su decimosexta edición en el Instituto Superior Técnico de Lisboa, Portugal, del 20 al 25 de julio. Es una de las escuelas de verano de aprendizaje automático más reconocidas de Europa y forma parte de la red europea ELLIS.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Jaziel A. Carballo Tadeo fue aceptado para participar en esta edición.&lt;/strong&gt; La escuela combina clases magistrales matutinas con laboratorios prácticos por la tarde, presentaciones de pósteres y charlas invitadas, cubriendo temas que van de los modelos lineales a los transformers, la causalidad y los modelos visión-lenguaje. Esta formación fortalece directamente la investigación del grupo en tecnologías del lenguaje para el maya yucateco.&lt;/p&gt;</description></item><item><title>LxMLS 2026: The 16th Lisbon Machine Learning School</title><link>https://giltia.github.io/event/lxmls-2026/</link><pubDate>Mon, 20 Jul 2026 09:00:00 +0000</pubDate><guid>https://giltia.github.io/event/lxmls-2026/</guid><description>&lt;p&gt;The &lt;a href="https://lxmls.github.io/2026/" target="_blank" rel="noopener"&gt;Lisbon Machine Learning School (LxMLS)&lt;/a&gt; holds its sixteenth edition in 2026 at Instituto Superior Técnico in Lisbon, Portugal, from July 20th to 25th. It is one of Europe&amp;rsquo;s most recognized machine learning summer schools and is part of the European ELLIS network.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Jaziel A. Carballo Tadeo was accepted to participate in this edition.&lt;/strong&gt; The school combines morning lectures with hands-on afternoon labs, poster presentations, and invited talks, covering topics ranging from linear models to transformers, causality, and vision-language models. This training directly strengthens the group&amp;rsquo;s research on language technologies for Yucatec Maya.&lt;/p&gt;</description></item><item><title>Whisper-LM con léxico maya: corrección de transcripciones ASR de maya yucateco usando Gemini y un diccionario como referencia</title><link>https://giltia.github.io/es/publication/whisper-lm-yua-lexicon-rag/</link><pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/publication/whisper-lm-yua-lexicon-rag/</guid><description>&lt;h2 id="qué-problema-resuelve-este-cuaderno"&gt;¿Qué problema resuelve este cuaderno?&lt;/h2&gt;
&lt;p&gt;Whisper, el sistema de reconocimiento automático de voz (ASR) de OpenAI, no incluye el maya yucateco entre sus lenguas soportadas. Al transcribir audio en maya, el modelo produce salidas muy ruidosas y con frecuencia &amp;ldquo;detecta&amp;rdquo; la lengua equivocada: en nuestros datos, segmentos en maya fueron etiquetados como español, inglés e incluso japonés. Este cuaderno explora una pregunta práctica: &lt;strong&gt;¿cuánto puede mejorar un modelo de lenguaje grande (LLM) la salida de Whisper si le damos un léxico maya como referencia?&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="arquitectura-del-pipeline"&gt;Arquitectura del pipeline&lt;/h2&gt;
&lt;p&gt;El flujo completo es:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Audio → Whisper (turbo) → Normalizador de reglas → Gemini + léxico → Normalizador → Evaluación (CER/WER)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="1-datos-de-entrada"&gt;1. Datos de entrada&lt;/h3&gt;
&lt;p&gt;Se parte de un CSV (&lt;code&gt;transcription_results41Turbo.csv&lt;/code&gt;) con 1,010 segmentos de audio ya transcritos por Whisper turbo. Cada fila incluye la transcripción de referencia hecha por humanos (&lt;code&gt;original_transcription&lt;/code&gt;), la predicción de Whisper (&lt;code&gt;whisper_prediction&lt;/code&gt;), la lengua detectada y métricas iniciales (chrF, WER, CER).&lt;/p&gt;
&lt;h3 id="2-léxico-maya-como-referencia-componente-rag"&gt;2. Léxico maya como referencia (componente RAG)&lt;/h3&gt;
&lt;p&gt;Se carga un diccionario maya–español en formato TSV (&lt;code&gt;yua_dictionary.tsv&lt;/code&gt;) con &lt;strong&gt;2,988 entradas&lt;/strong&gt;. Una clase &lt;code&gt;YuaLexicon&lt;/code&gt; ofrece búsqueda exacta y búsqueda difusa (&lt;em&gt;fuzzy matching&lt;/em&gt; con &lt;code&gt;difflib.get_close_matches&lt;/code&gt;), lo que permite encontrar la palabra maya válida más cercana a un token mal transcrito.&lt;/p&gt;
&lt;h3 id="3-normalizador-ortográfico-basado-en-reglas"&gt;3. Normalizador ortográfico basado en reglas&lt;/h3&gt;
&lt;p&gt;Antes y después de pasar por el LLM, el texto se normaliza con reglas lingüísticas del maya yucateco tomadas de una gramática de referencia:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Reglas morfológicas de raíz&lt;/strong&gt;: por ejemplo, &lt;em&gt;bin&lt;/em&gt; (ir) con raíz irregular &lt;em&gt;xi&amp;rsquo;&lt;/em&gt; en futuro indefinido intransitivo, o posicionales como &lt;em&gt;chil&lt;/em&gt;, &lt;em&gt;kul&lt;/em&gt; y &lt;em&gt;wa&amp;rsquo;al&lt;/em&gt; que pierden la &lt;em&gt;l&lt;/em&gt; antes de &lt;em&gt;-tal&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Patrones regulares &amp;ldquo;seguros&amp;rdquo;&lt;/strong&gt;: &lt;em&gt;bins&lt;/em&gt; → &lt;em&gt;bis&lt;/em&gt;, &lt;em&gt;taals&lt;/em&gt; → &lt;em&gt;taas&lt;/em&gt;, y verbos que cambian la &lt;em&gt;b&lt;/em&gt; final por estructura con saltillo (&lt;em&gt;jáalk&amp;rsquo;ab&lt;/em&gt; → &lt;em&gt;jáalk&amp;rsquo;a&amp;rsquo;a&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Verificación contra el léxico&lt;/strong&gt;: si la forma normalizada existe en el diccionario, se conserva; si la original era válida, se respeta.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="4-corrección-con-gemini-whisper-lm"&gt;4. Corrección con Gemini (Whisper-LM)&lt;/h3&gt;
&lt;p&gt;Cada transcripción normalizada se envía a &lt;strong&gt;Gemini 2.0 Flash&lt;/strong&gt; con un prompt que incluye hasta 500 palabras del léxico maya como referencia ortográfica. Las instrucciones clave del prompt son: corregir respetando la gramática y ortografía estándar del maya yucateco, &lt;strong&gt;no traducir al español&lt;/strong&gt;, elegir la opción más probable según el léxico y devolver únicamente el texto corregido.&lt;/p&gt;
&lt;h3 id="5-evaluación"&gt;5. Evaluación&lt;/h3&gt;
&lt;p&gt;Se calcula la tasa de error por carácter (CER) y por palabra (WER) con &lt;code&gt;editdistance&lt;/code&gt;, comparando contra la transcripción humana de referencia:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Métrica&lt;/th&gt;
&lt;th&gt;Whisper turbo&lt;/th&gt;
&lt;th&gt;Whisper-LM (este pipeline)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CER medio&lt;/td&gt;
&lt;td&gt;1.0743&lt;/td&gt;
&lt;td&gt;1.0725&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WER medio&lt;/td&gt;
&lt;td&gt;1.1796&lt;/td&gt;
&lt;td&gt;1.1692&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="qué-aprendimos"&gt;¿Qué aprendimos?&lt;/h2&gt;
&lt;p&gt;La mejora es &lt;strong&gt;marginal&lt;/strong&gt;: el posprocesamiento con LLM y léxico recupera algo de forma ortográfica, pero no puede reconstruir información que el modelo acústico nunca capturó. Con valores de CER/WER superiores a 1.0, la salida de Whisper está tan alejada de la referencia que la corrección textual tiene poco material con el que trabajar. La conclusión práctica es que, para el maya yucateco, &lt;strong&gt;el cuello de botella está en el modelo acústico&lt;/strong&gt;: se necesita ajuste fino (&lt;em&gt;fine-tuning&lt;/em&gt;) de Whisper con datos de habla maya, y no solo corrección posterior.&lt;/p&gt;
&lt;h2 id="próximos-pasos"&gt;Próximos pasos&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Ajuste fino de Whisper con audio transcrito en maya yucateco.&lt;/li&gt;
&lt;li&gt;Ampliar las reglas ortográficas (&lt;code&gt;ORTHO_RULES&lt;/code&gt;) con los errores recurrentes observados en las transcripciones.&lt;/li&gt;
&lt;li&gt;Ampliar el léxico y experimentar con recuperación selectiva (enviar al prompt solo las entradas relevantes para cada segmento, en lugar de una lista fija).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="acceso-al-cuaderno"&gt;Acceso al cuaderno&lt;/h2&gt;
&lt;p&gt;Si te interesa el cuaderno de Colab, puedes ponerte en contacto con el autor en &lt;a href="mailto:jaziel.carballo@gmail.com"&gt;jaziel.carballo@gmail.com&lt;/a&gt;.&lt;/p&gt;</description></item><item><title>Whisper-LM with a Maya Lexicon: Correcting Yucatec Maya ASR Transcriptions Using Gemini and a Dictionary as Reference</title><link>https://giltia.github.io/publication/whisper-lm-yua-lexicon-rag/</link><pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/publication/whisper-lm-yua-lexicon-rag/</guid><description>&lt;h2 id="what-problem-does-this-notebook-address"&gt;What problem does this notebook address?&lt;/h2&gt;
&lt;p&gt;Whisper, OpenAI&amp;rsquo;s automatic speech recognition (ASR) system, does not include Yucatec Maya among its supported languages. When transcribing Maya audio, the model produces very noisy output and frequently &amp;ldquo;detects&amp;rdquo; the wrong language: in our data, Maya segments were labeled as Spanish, English, and even Japanese. This notebook explores a practical question: &lt;strong&gt;how much can a large language model (LLM) improve Whisper&amp;rsquo;s output if we give it a Maya lexicon as reference?&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="pipeline-architecture"&gt;Pipeline architecture&lt;/h2&gt;
&lt;p&gt;The full flow is:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Audio → Whisper (turbo) → Rule-based normalizer → Gemini + lexicon → Normalizer → Evaluation (CER/WER)
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="1-input-data"&gt;1. Input data&lt;/h3&gt;
&lt;p&gt;The starting point is a CSV file (&lt;code&gt;transcription_results41Turbo.csv&lt;/code&gt;) with 1,010 audio segments already transcribed by Whisper turbo. Each row includes the human reference transcription (&lt;code&gt;original_transcription&lt;/code&gt;), Whisper&amp;rsquo;s prediction (&lt;code&gt;whisper_prediction&lt;/code&gt;), the detected language, and initial metrics (chrF, WER, CER).&lt;/p&gt;
&lt;h3 id="2-maya-lexicon-as-reference-rag-component"&gt;2. Maya lexicon as reference (RAG component)&lt;/h3&gt;
&lt;p&gt;A Maya–Spanish dictionary in TSV format (&lt;code&gt;yua_dictionary.tsv&lt;/code&gt;) with &lt;strong&gt;2,988 entries&lt;/strong&gt; is loaded. A &lt;code&gt;YuaLexicon&lt;/code&gt; class provides exact lookup and fuzzy matching (via &lt;code&gt;difflib.get_close_matches&lt;/code&gt;), making it possible to find the closest valid Maya word for a mis-transcribed token.&lt;/p&gt;
&lt;h3 id="3-rule-based-orthographic-normalizer"&gt;3. Rule-based orthographic normalizer&lt;/h3&gt;
&lt;p&gt;Before and after the LLM pass, the text is normalized with Yucatec Maya linguistic rules taken from a reference grammar:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Root morphology rules&lt;/strong&gt;: for example, &lt;em&gt;bin&lt;/em&gt; (to go) with the irregular root &lt;em&gt;xi&amp;rsquo;&lt;/em&gt; in the intransitive indefinite future, or positionals such as &lt;em&gt;chil&lt;/em&gt;, &lt;em&gt;kul&lt;/em&gt;, and &lt;em&gt;wa&amp;rsquo;al&lt;/em&gt; that lose the &lt;em&gt;l&lt;/em&gt; before &lt;em&gt;-tal&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Safe&amp;rdquo; regex patterns&lt;/strong&gt;: &lt;em&gt;bins&lt;/em&gt; → &lt;em&gt;bis&lt;/em&gt;, &lt;em&gt;taals&lt;/em&gt; → &lt;em&gt;taas&lt;/em&gt;, and verbs whose final &lt;em&gt;b&lt;/em&gt; changes to a glottalized structure (&lt;em&gt;jáalk&amp;rsquo;ab&lt;/em&gt; → &lt;em&gt;jáalk&amp;rsquo;a&amp;rsquo;a&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Lexicon verification&lt;/strong&gt;: if the normalized form exists in the dictionary it is kept; if the original form was already valid, it is preserved.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="4-correction-with-gemini-whisper-lm"&gt;4. Correction with Gemini (Whisper-LM)&lt;/h3&gt;
&lt;p&gt;Each normalized transcription is sent to &lt;strong&gt;Gemini 2.0 Flash&lt;/strong&gt; with a prompt that includes up to 500 words from the Maya lexicon as an orthographic reference. The key prompt instructions are: correct the text following standard Yucatec Maya grammar and orthography, &lt;strong&gt;do not translate into Spanish&lt;/strong&gt;, choose the most likely option according to the lexicon, and return only the corrected text.&lt;/p&gt;
&lt;h3 id="5-evaluation"&gt;5. Evaluation&lt;/h3&gt;
&lt;p&gt;Character error rate (CER) and word error rate (WER) are computed with &lt;code&gt;editdistance&lt;/code&gt; against the human reference transcription:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Whisper turbo&lt;/th&gt;
&lt;th&gt;Whisper-LM (this pipeline)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mean CER&lt;/td&gt;
&lt;td&gt;1.0743&lt;/td&gt;
&lt;td&gt;1.0725&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mean WER&lt;/td&gt;
&lt;td&gt;1.1796&lt;/td&gt;
&lt;td&gt;1.1692&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="what-did-we-learn"&gt;What did we learn?&lt;/h2&gt;
&lt;p&gt;The improvement is &lt;strong&gt;marginal&lt;/strong&gt;: LLM post-processing with a lexicon recovers some orthographic form, but it cannot reconstruct information the acoustic model never captured. With CER/WER values above 1.0, Whisper&amp;rsquo;s output is so far from the reference that text-level correction has little material to work with. The practical takeaway is that, for Yucatec Maya, &lt;strong&gt;the bottleneck is the acoustic model&lt;/strong&gt;: what is needed is fine-tuning Whisper on Maya speech data, not just downstream correction.&lt;/p&gt;
&lt;h2 id="next-steps"&gt;Next steps&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Fine-tune Whisper with transcribed Yucatec Maya audio.&lt;/li&gt;
&lt;li&gt;Extend the orthographic rules (&lt;code&gt;ORTHO_RULES&lt;/code&gt;) with recurrent errors observed in the transcriptions.&lt;/li&gt;
&lt;li&gt;Grow the lexicon and experiment with selective retrieval (sending only the entries relevant to each segment to the prompt, instead of a fixed list).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="access-to-the-notebook"&gt;Access to the notebook&lt;/h2&gt;
&lt;p&gt;If you are interested in the Colab notebook, you can contact the author at &lt;a href="mailto:jaziel.carballo@gmail.com"&gt;jaziel.carballo@gmail.com&lt;/a&gt;.&lt;/p&gt;</description></item><item><title>CentroGeo: introduces the Maya–Spanish parallel corpus YUA-ES-CCC</title><link>https://giltia.github.io/post/centrogeo-yua-es-corpus/</link><pubDate>Wed, 01 Jul 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/post/centrogeo-yua-es-corpus/</guid><description>&lt;p&gt;&lt;strong&gt;CentroGeo Centro Público de Investigación&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;We present the YUA-ES Communicative Contexts Corpus (YUA-ES-CCC), an open collection of 14,332 phrase pairs in Yucatec Maya and Spanish, organized into 31 everyday contexts: greetings, transportation, everyday conversation, family, food, and more.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;More than 14,000 phrases in Yucatec Maya and Spanish, available as open data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Maya–Spanish parallel corpus (YUA-ES-CCC)&lt;/strong&gt;
An open resource that drives research, education, and the development of language technologies for Indigenous languages.&lt;/p&gt;
&lt;p&gt;Learn more at: cgeo.mx/#d35&lt;/p&gt;
&lt;p&gt;Project developed in collaboration with Sedeculta Yucatán (Secretaría de Ciencias, Humanidades, Tecnología e Innovación).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The corpus is described in the publication &lt;a href="https://giltia.github.io/publication/yua-es-corpus/"&gt;&lt;em&gt;The YUA-ES Communicative Contexts Corpus&lt;/em&gt;&lt;/a&gt;, available as a preprint on Research Square.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Post by CentroGeo Centro Público de Investigación on Facebook, July 1, 2026. Translated from the original Spanish for this archive.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>CentroGeo: presenta el corpus paralelo maya–español YUA-ES-CCC</title><link>https://giltia.github.io/es/post/centrogeo-yua-es-corpus/</link><pubDate>Wed, 01 Jul 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/post/centrogeo-yua-es-corpus/</guid><description>&lt;p&gt;&lt;strong&gt;CentroGeo Centro Público de Investigación&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Presentamos el YUA-ES Communicative Contexts Corpus (YUA-ES-CCC), una colección abierta de 14,332 pares de frases en maya yucateco y español, organizadas en 31 contextos cotidianos: saludos, transporte, conversación cotidiana, familia, comida y más.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Más de 14 mil frases en maya yucateco y español, disponibles como datos abiertos.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Corpus paralelo maya–español (YUA-ES-CCC)&lt;/strong&gt;
Un recurso abierto que impulsa la investigación, la educación y la generación de tecnologías lingüísticas para las lenguas indígenas.&lt;/p&gt;
&lt;p&gt;Conoce más en: cgeo.mx/#d35&lt;/p&gt;
&lt;p&gt;Proyecto desarrollado en colaboración con Sedeculta Yucatán (Secretaría de Ciencias, Humanidades, Tecnología e Innovación).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;El corpus está descrito en la publicación &lt;a href="https://giltia.github.io/es/publication/yua-es-corpus/"&gt;&lt;em&gt;The YUA-ES Communicative Contexts Corpus&lt;/em&gt;&lt;/a&gt;, disponible como preprint en Research Square.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Publicación de CentroGeo Centro Público de Investigación en Facebook, 1 de julio de 2026.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>The YUA-ES Communicative Contexts Corpus: An Open Parallel Dataset of Everyday Yucatec Maya–Spanish Phrases</title><link>https://giltia.github.io/es/publication/yua-es-corpus/</link><pubDate>Wed, 17 Jun 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/publication/yua-es-corpus/</guid><description/></item><item><title>The YUA-ES Communicative Contexts Corpus: An Open Parallel Dataset of Everyday Yucatec Maya–Spanish Phrases</title><link>https://giltia.github.io/publication/yua-es-corpus/</link><pubDate>Wed, 17 Jun 2026 00:00:00 +0000</pubDate><guid>https://giltia.github.io/publication/yua-es-corpus/</guid><description/></item><item><title>Generating a Culturally and Linguistically Adapted Word Similarity Benchmark for Yucatec Maya</title><link>https://giltia.github.io/es/publication/maya-word-similarity-benchmark/</link><pubDate>Thu, 25 Sep 2025 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/publication/maya-word-similarity-benchmark/</guid><description/></item><item><title>Generating a Culturally and Linguistically Adapted Word Similarity Benchmark for Yucatec Maya</title><link>https://giltia.github.io/publication/maya-word-similarity-benchmark/</link><pubDate>Thu, 25 Sep 2025 00:00:00 +0000</pubDate><guid>https://giltia.github.io/publication/maya-word-similarity-benchmark/</guid><description/></item><item><title>Findings of the AmericasNLP 2024 Shared Task on the Creation of Educational Materials for Indigenous Languages</title><link>https://giltia.github.io/es/publication/americasnlp-2024-findings/</link><pubDate>Fri, 21 Jun 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/publication/americasnlp-2024-findings/</guid><description/></item><item><title>Findings of the AmericasNLP 2024 Shared Task on the Creation of Educational Materials for Indigenous Languages</title><link>https://giltia.github.io/publication/americasnlp-2024-findings/</link><pubDate>Fri, 21 Jun 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/publication/americasnlp-2024-findings/</guid><description/></item><item><title>Mayasoundex: A Phonetically Grounded Algorithm for Information Retrieval in the Maya Language</title><link>https://giltia.github.io/es/publication/mayasoundex/</link><pubDate>Sun, 16 Jun 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/publication/mayasoundex/</guid><description/></item><item><title>Mayasoundex: A Phonetically Grounded Algorithm for Information Retrieval in the Maya Language</title><link>https://giltia.github.io/publication/mayasoundex/</link><pubDate>Sun, 16 Jun 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/publication/mayasoundex/</guid><description/></item><item><title>GeoInt Difusión: It might be possible to browse the web in Maya</title><link>https://giltia.github.io/post/geoint-taantsil-navegar-maya-web/</link><pubDate>Mon, 01 Apr 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/post/geoint-taantsil-navegar-maya-web/</guid><description>&lt;p&gt;&lt;strong&gt;Difusión | News&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="it-might-be-possible-to-browse-the-web-in-maya"&gt;It might be possible to browse the web in Maya&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;Iván Canul Ek&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The article introduces the T&amp;rsquo;aantsil project, a linguistic corpus platform for Yucatec Maya led by Alejandro Molina-Villegas, a Conahcyt researcher based at CentroGeo.&lt;/p&gt;
&lt;p&gt;The initiative seeks to bring the Maya language into the digital space, allowing its speakers to access technologies already available in other languages. T&amp;rsquo;aantsil compiles multimedia material from Maya-speaking communities through interviews, videos, and audio recordings that are transcribed in Maya, translated into Spanish, and then into English using artificial intelligence.&lt;/p&gt;
&lt;p&gt;Molina-Villegas notes that &amp;ldquo;the Maya language lags behind in the digital space compared with other, better-represented languages,&amp;rdquo; and points out that the platform helps preserve the language amid its risk of disappearing, as the number of speakers continues to decline.&lt;/p&gt;
&lt;p&gt;The platform, available at taantsil.com.mx and developed with César Can Canul of Uady, allows users to search for Maya words and see interview excerpts, transcriptions, and translations, including regional variants.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Article by Iván Canul Ek, published in GeoInt Difusión (CentroGeo) in April 2024. Translated from the original Spanish for this archive.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>GeoInt Difusión: Se podría navegar en maya en la web</title><link>https://giltia.github.io/es/post/geoint-taantsil-navegar-maya-web/</link><pubDate>Mon, 01 Apr 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/post/geoint-taantsil-navegar-maya-web/</guid><description>&lt;p&gt;&lt;strong&gt;Difusión | Noticias&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="se-podría-navegar-en-maya-en-la-web"&gt;Se podría navegar en maya en la web&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;Iván Canul Ek&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;El artículo presenta el proyecto T&amp;rsquo;aantsil, una plataforma de corpus lingüístico para el maya yucateco dirigida por Alejandro Molina-Villegas, investigador del Conahcyt adscrito a CentroGeo.&lt;/p&gt;
&lt;p&gt;La iniciativa busca integrar la lengua maya al espacio digital, permitiendo que sus hablantes accedan a tecnologías disponibles en otros idiomas. T&amp;rsquo;aantsil compila material multimedia de comunidades mayahablantes mediante entrevistas, videos y audios que se transcriben en maya, traducen al español y posteriormente al inglés mediante inteligencia artificial.&lt;/p&gt;
&lt;p&gt;Molina-Villegas señala que &amp;ldquo;la lengua maya está rezagada en el espacio digital con respecto a otras lenguas mejor representadas&amp;rdquo; y destaca que la plataforma preserva la lengua ante su riesgo de desaparición, pues hay cada vez menos hablantes.&lt;/p&gt;
&lt;p&gt;La plataforma, disponible en taantsil.com.mx y desarrollada con César Can Canul, de la Uady, permite búsquedas de palabras en maya con fragmentos de entrevistas, transcripciones y traducciones, incluyendo variantes regionales.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Nota de Iván Canul Ek, publicada en GeoInt Difusión (CentroGeo) en abril de 2024.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Contact</title><link>https://giltia.github.io/contact/</link><pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/contact/</guid><description/></item><item><title>Contacto</title><link>https://giltia.github.io/es/contact/</link><pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/contact/</guid><description/></item><item><title>Equipo</title><link>https://giltia.github.io/es/people/</link><pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/people/</guid><description/></item><item><title>People</title><link>https://giltia.github.io/people/</link><pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate><guid>https://giltia.github.io/people/</guid><description/></item><item><title>Tour</title><link>https://giltia.github.io/es/tour/</link><pubDate>Mon, 24 Oct 2022 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/tour/</guid><description/></item><item><title>Tour</title><link>https://giltia.github.io/tour/</link><pubDate>Mon, 24 Oct 2022 00:00:00 +0000</pubDate><guid>https://giltia.github.io/tour/</guid><description/></item><item><title/><link>https://giltia.github.io/admin/config.yml</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://giltia.github.io/admin/config.yml</guid><description/></item><item><title/><link>https://giltia.github.io/es/admin/config.yml</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://giltia.github.io/es/admin/config.yml</guid><description/></item></channel></rss>