CoMMA (Corpus of Multilingual Medieval Archives)
CoMMA is a large-scale corpus of medieval manuscripts produced through automatic text recognition. The corpus contains around 3.3b tokens drawn from more than 32,700 digitized manuscripts in Latin and Old French, harvested via IIIF. Unlike other resources, it is made of raw, non-normalized text enriched with layout analysis in various formats.
Accès libre (cc-by-4.0)
https://comma.inria.fr/homepage /
https://huggingface.co/datasets/comma-project/comma-jsonl
– Aucune version pour une plateforme spécifique
Référence
Thibault Clérice, Simon Gabay, Malamatenia Vlachou-Efstathiou, Ariane Pinche, Benoît Sagot. “CoMMA, a Large-scale Corpus of Multilingual Medieval Archives”. LREC 2026 – 15th edition of the Language Resources and Evaluation Conference, May 2026, Palma de Mallorca, Spain. ⟨hal-05299220v2⟩.
