Your browser does not allow JavaScript!
JavaScript is necessary for the proper functioning of this website. Please enable JavaScript or use a modern browser.
Repository of the University of Ljubljana
Open Science Slovenia
Open Science
DiKUL
slv
|
eng
Search
Advanced
New in RUL
About RUL
In numbers
Help
Sign in
Details
Evaluating robustness of LLMs in question answering on multilingual noisy OCR data
ID
Piryani, Bhawna
(
Author
),
ID
Mozafari, Jamshid
(
Author
),
ID
Abdallah, Abdelrahman
(
Author
),
ID
Doucet, Antoine
(
Author
),
ID
Jatowt, Adam
(
Author
)
URL - Source URL, Visit
https://doi.org/10.1145/3746252.3761295
PDF - Presentation file,
Download
(1,38 MB)
MD5: A9B7F26E6F89C804437345EC8B7760D1
Image galllery
Abstract
Optical Character Recognition (OCR) plays a crucial role in digitizing historical and multilingual documents, yet OCR errors - imperfect extraction of text, including character insertion, deletion, and substitution can significantly impact downstream tasks like question-answering (QA). In this work, we conduct a comprehensive analysis of how OCR-induced noise affects the performance of Multilingual QA Systems. To support this analysis, we introduce a multilingual QA dataset MultiOCR-QA, comprising 50K question-answer pairs across three languages, English, French, and German. The dataset is curated from OCR-ed historical documents, which include different levels and types of OCR noise. We then evaluate how different state-of-the-art Large Language Models (LLMs) perform under different error conditions, focusing on three major OCR error types. Our findings show that QA systems are highly prone to OCR-induced errors and perform poorly on noisy OCR text. By comparing model performance on clean versus noisy texts, we provide insights into the limitations of current approaches and emphasize the need for more noise-resilient QA systems in historical digitization contexts.
Language:
English
Keywords:
multilingual QA
,
OCR text
,
large language models
Work type:
Other
Typology:
1.08 - Published Scientific Conference Contribution
Organization:
FRI - Faculty of Computer and Information Science
Publication status:
Published
Publication version:
Version of Record
Year:
2025
Number of pages:
Str. 2366-2376
PID:
20.500.12556/RUL-181096
UDC:
004.85:004.352.242:81'322
DOI:
10.1145/3746252.3761295
COBISS.SI-ID:
272786691
Publication date in RUL:
25.03.2026
Views:
384
Downloads:
183
Metadata:
Cite this work
Plain text
BibTeX
EndNote XML
EndNote/Refer
RIS
ABNT
ACM Ref
AMA
APA
Chicago 17th Author-Date
Harvard
IEEE
ISO 690
MLA
Vancouver
:
Copy citation
Share:
Record is a part of a monograph
Title:
CIKM ’25 : proceedings of the 34th ACM International Conference on Information and Knowledge Management
Place of publishing:
New York (NY)
Publisher:
The Association for Computing Machinery
ISBN:
979-8-4007-2040-6
COBISS.SI-ID:
272764675
Licences
License:
CC BY 4.0, Creative Commons Attribution 4.0 International
Link:
http://creativecommons.org/licenses/by/4.0/
Description:
This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.
Secondary language
Language:
Slovenian
Keywords:
večjezično zagotavljanje kakovosti
,
optično prepoznavanje besedila
,
veliki jezikovni modeli
Projects
Funder:
EC - European Commission
Project number:
101186647
Name:
Centre of Excellence in Artificial Intelligence for Digital Humanities
Acronym:
AI4DH
Similar documents
Similar works from RUL:
Similar works from other Slovenian collections:
Back