Researchers at Inria have reached a major milestone in digital history. By creating a tool called CoMMA, they successfully transcribed over 32,000 medieval manuscripts in just four months. This task usually consumes the entire career of a human researcher, but a new approach focused on visual character recognition has changed the math.

The challenge with ancient documents is not just the handwriting. Medieval texts are filled with complex abbreviations, and languages like Old French lack standardized spelling. While popular language models often hallucinate when faced with this lack of regularity, the team at Inria bypassed linguistic prediction entirely. They instead trained their software to recognize individual shapes. By treating an accent or a flourish as a unique character, the system maintains accuracy without guessing meaning.

This project relied on the foundation of the CATMuS initiative, which spent years creating a massive, consistent dataset of hand-transcribed lines. The researchers avoided correcting mistakes or resolving abbreviations, choosing instead to preserve the raw, historical integrity of the source material. This method keeps the model grounded in what is actually on the page.

Because of this work, scholars now have access to more than three billion words of Latin and Old French. The volume of available Old French text alone has increased forty-fold. Accessing these digitized archives allows historians to study language and society on a scale that was previously impossible. The CoMMA platform remains free and public for all users to explore.