Details

Carniolan Provincial Assembly : corpus improvements and enhancements
ID Pretnar Žagar, Ajda (Author), ID Pahor de Maiti, Kristina (Author)

.pdfPDF - Presentation file, Download (1,25 MB)
MD5: 51E40C830BDCD61D747CB03EB2DF8E00
URLURL - Source URL, Visit https://journals.uio.no/dhnbpub/article/view/13202 This link opens in a new window

Abstract
Historical parliamentary corpora offer crucial evidence for studying political discourse over time, yet their usability is often limited by poor OCR quality and incomplete metadata. This paper presents the enhancement of the Kranjska 1.0 corpus, a collection of Carniolan Provincial Assembly proceedings (1861–1913) in Slovenian and German, through a two-phase process aimed at improving textual accuracy and enriching speaker metadata. First, we conducted a manual correction campaign on a representative sample of transcripts, involving trained historians proficient in Gothic script and 19th-century politics. The corrections addressed both structural and textual errors in TEI-encoded XML files, providing a gold-standard dataset for future model training. Error analysis revealed recurring OCR issues, including segmentation problems, misattributed speakers, and systematic character-level noise. Second, we harmonised and expanded speaker metadata using multiple historical sources to unify name variants, resolve ambiguities, and document parliamentary terms, factions, and attendance. The resulting metadata enhance corpus usability and interpretability. This work lays the foundation for the next project phase, which explores the automatic correction of transcripts and metadata using Multimodal Large Language Models (MLLMs). By combining historical expertise with computational methods, we contribute to more accurate processing of historical texts and promote transparency and reusability in digital humanities research.

Language:English
Keywords:historical parliamentary proceedings, OCR correction, error analysis, metadata enrichment
Work type:Other
Typology:1.08 - Published Scientific Conference Contribution
Organization:FRI - Faculty of Computer and Information Science
FF - Faculty of Arts
Publication status:Published
Publication version:Version of Record
Year:2026
Number of pages:Str. 1-10
PID:20.500.12556/RUL-180816 This link opens in a new window
UDC:004.89:328(497.12)”1861/1913”
ISSN on article:2704-1441
DOI:10.5617/dhnbpub.13202 This link opens in a new window
COBISS.SI-ID:271706883 This link opens in a new window
Publication date in RUL:17.03.2026
Views:354
Downloads:170
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Record is a part of a proceedings

Title:Lost in abundance
COBISS.SI-ID:271613699 This link opens in a new window

Record is a part of a journal

Title:Digital humanities in the Nordic and Baltic countries publications : DHNB
Publisher:University of Oslo library
ISSN:2704-1441
COBISS.SI-ID:228389891 This link opens in a new window

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Keywords:zgodovinski parlamentarni zapisniki, popravki napak OCR, analiza napak, obogatitev metapodatkov

Projects

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:P6-0436-2022
Name:Digitalna humanistika: viri, orodja in metode

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:GC-0002-2024
Name:Veliki jezikovni modeli za digitalno humanistiko

Funder:EC - European Commission
Funding programme:HE
Project number:101186647
Name:Centre of Excellence in Artificial Intelligence for Digital Humanities
Acronym:AI4DH

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back