Details

Evaluating robustness of LLMs in question answering on multilingual noisy OCR data
ID Piryani, Bhawna (Author), ID Mozafari, Jamshid (Author), ID Abdallah, Abdelrahman (Author), ID Doucet, Antoine (Author), ID Jatowt, Adam (Author)

URLURL - Source URL, Visit https://doi.org/10.1145/3746252.3761295 This link opens in a new window
.pdfPDF - Presentation file, Download (1,38 MB)
MD5: A9B7F26E6F89C804437345EC8B7760D1

Abstract
Optical Character Recognition (OCR) plays a crucial role in digitizing historical and multilingual documents, yet OCR errors - imperfect extraction of text, including character insertion, deletion, and substitution can significantly impact downstream tasks like question-answering (QA). In this work, we conduct a comprehensive analysis of how OCR-induced noise affects the performance of Multilingual QA Systems. To support this analysis, we introduce a multilingual QA dataset MultiOCR-QA, comprising 50K question-answer pairs across three languages, English, French, and German. The dataset is curated from OCR-ed historical documents, which include different levels and types of OCR noise. We then evaluate how different state-of-the-art Large Language Models (LLMs) perform under different error conditions, focusing on three major OCR error types. Our findings show that QA systems are highly prone to OCR-induced errors and perform poorly on noisy OCR text. By comparing model performance on clean versus noisy texts, we provide insights into the limitations of current approaches and emphasize the need for more noise-resilient QA systems in historical digitization contexts.

Language:English
Keywords:multilingual QA, OCR text, large language models
Work type:Other
Typology:1.08 - Published Scientific Conference Contribution
Organization:FRI - Faculty of Computer and Information Science
Publication status:Published
Publication version:Version of Record
Year:2025
Number of pages:Str. 2366-2376
PID:20.500.12556/RUL-181096 This link opens in a new window
UDC:004.85:004.352.242:81'322
DOI:10.1145/3746252.3761295 This link opens in a new window
COBISS.SI-ID:272786691 This link opens in a new window
Publication date in RUL:25.03.2026
Views:384
Downloads:183
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Record is a part of a monograph

Title:CIKM ’25 : proceedings of the 34th ACM International Conference on Information and Knowledge Management
Place of publishing:New York (NY)
Publisher:The Association for Computing Machinery
ISBN:979-8-4007-2040-6
COBISS.SI-ID:272764675 This link opens in a new window

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Keywords:večjezično zagotavljanje kakovosti, optično prepoznavanje besedila, veliki jezikovni modeli

Projects

Funder:EC - European Commission
Project number:101186647
Name:Centre of Excellence in Artificial Intelligence for Digital Humanities
Acronym:AI4DH

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back