Your browser does not allow JavaScript!
JavaScript is necessary for the proper functioning of this website. Please enable JavaScript or use a modern browser.
Repository of the University of Ljubljana
Open Science Slovenia
Open Science
DiKUL
slv
|
eng
Search
Advanced
New in RUL
About RUL
In numbers
Help
Sign in
Details
Assessing reliability of BERT-based models on question answering tasks
ID
Yadav, Pooja
(
Author
),
ID
Harjule, Priyanka
(
Author
),
ID
Agarwal, Basant
(
Author
),
ID
Robnik Šikonja, Marko
(
Author
)
PDF - Presentation file,
Download
(3,84 MB)
MD5: 9CAEE05F93DFB84F202CAB2319D9A59C
URL - Source URL, Visit
https://www.tandfonline.com/doi/full/10.1080/0952813X.2026.2716084
Image galllery
Abstract
Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suitable for practical applications. Recent advancements in natural language processing (NLP), particularly those based on transformer architectures, have significantly accelerated progress across various NLP tasks. This study focuses on the reliability of transformer-based question answering (QA) models, specifically BERT models and its variants (RoBERTa, ALBERT, DistilBERT). These encoder-only pretrained transformers have demonstrated remarkable accuracy in QA tasks that can be treated as classification tasks. However, their reliability remains underexplored. This study evaluates the reliability of four BERT-based models by assessing response stability under two conditions: (1) internal model variations induced via Monte Carlo Dropout (MCD) and (2) input perturbations through paraphrasing. Using the SQuAD and QuAC datasets, we investigate how dropout rates affect prediction consistency and whether lexical changes impact answer stability. Our findings reveal that RoBERTa maintains higher reliability, whereas AlBERT and DistilBERT exhibit significant inconsistencies. Statistical analyses confirm that enabling MCD during prediction does not disrupt inference dynamics, validating its effectiveness as a reliability metric. These findings underscore the importance of evaluating both accuracy and stability in QA models to ensure stability in real-world applications.
Language:
English
Keywords:
natural language processing
,
large language models
,
reliability estimation
,
BERT models
,
question answering
Work type:
Article
Typology:
1.01 - Original Scientific Article
Organization:
FRI - Faculty of Computer and Information Science
Publication status:
Published
Publication version:
Version of Record
Year:
2026
Number of pages:
21 str.
Numbering:
Vol. , no.
PID:
20.500.12556/RUL-186056
UDC:
004.85:004.912:81'322
ISSN on article:
0952-813X
DOI:
10.1080/0952813X.2026.2716084
COBISS.SI-ID:
288277507
Publication date in RUL:
26.08.2026
Views:
109
Downloads:
39
Metadata:
Cite this work
Plain text
BibTeX
EndNote XML
EndNote/Refer
RIS
ABNT
ACM Ref
AMA
APA
Chicago 17th Author-Date
Harvard
IEEE
ISO 690
MLA
Vancouver
:
Copy citation
Share:
Record is a part of a journal
Title:
Journal of experimental & theoretical artificial intelligence
Shortened title:
J. exp. theor. artif. intell.
Publisher:
Taylor & Francis
ISSN:
0952-813X
COBISS.SI-ID:
15368197
Licences
License:
CC BY 4.0, Creative Commons Attribution 4.0 International
Link:
http://creativecommons.org/licenses/by/4.0/
Description:
This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.
Secondary language
Language:
Slovenian
Keywords:
obdelava naravnega jezika
,
veliki jezikovni modeli
,
ocenjevanje zanesljivosti
,
modeli BERT
,
odgovarjanje na vprašanja
Projects
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
GC-0002-2024
Name:
Veliki jezikovni modeli za digitalno humanistiko
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
P6-0411-2019
Name:
Jezikovni viri in tehnologije za slovenski jezik
Funder:
EC - European Commission
Project number:
101186647
Name:
Centre of Excellence in Artificial Intelligence for Digital Humanities
Acronym:
AI4DH
Similar documents
Similar works from RUL:
Similar works from other Slovenian collections:
Back