Details

Prepoznavanje anomalij v avtomatsko ekstrahiranih grafih iz spleta
ID Safić, Sanil (Author), ID Žitnik, Slavko (Mentor) More about this mentor... This link opens in a new window

.pdfPDF - Presentation file, Download (1,00 MB)
MD5: 279E7A3C2991A3CF0D34A91B278CB63D

Abstract
V nalogi je predstavljen sistem za zaznavanje anomalij v avtomatsko ekstrahiranih grafih iz spleta, zgrajen na grafni bazi Neo4j in pravilih v jeziku Cypher. Sistem prepoznava strukturne, atributne in časovne nepravilnosti, kot so nenavadne lastniške strukture, nelogične investicije ter neskladja v datumih dogodkov, rezultate pa semantično preveri z velikim jezikovnim modelom (LLM). Evalvacija na grafu s približno 40 milijoni vozlišč pokaže visoko natančnost pri izbranih pravilih, pomemben vpliv materializiranih povezav na čas izvajanja poizvedb ter zmanjšanje obremenitve ročne validacije (QA). Pristop združuje razložljivost pravil s semantično analizo LLM in predstavlja korak k modularnemu, samoučečemu sistemu za zagotavljanje kakovosti podatkov.

Language:Slovenian
Keywords:Zaznavanje anomalij, grafne podatkovne baze, Neo4j, avtomatska ekstrakcija podatkov
Work type:Master's thesis/paper
Typology:2.09 - Master's Thesis
Organization:FRI - Faculty of Computer and Information Science
Year:2026
PID:20.500.12556/RUL-178466 This link opens in a new window
COBISS.SI-ID:268544259 This link opens in a new window
Publication date in RUL:28.01.2026
Views:372
Downloads:151
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Anomaly detection in automatically extracted graphs from the web
Abstract:
This thesis presents a system for detecting anomalies in automatically extracted graphs from the web, built on the Neo4j graph database and Cypherbased rules. The system identifies structural, attribute and temporal irregularities— such as unusual ownership structures, illogical investments and inconsistencies in event dates—and semantically validates the results with a Large Language Model (LLM). Evaluation on a graph with approximately 40 million nodes shows high precision for selected rules, a significant impact of materialized relationships on query runtimes, and a reduction of manual quality assurance (QA) workload. The approach combines the interpretability of rules with LLM-based semantic analysis and represents a step towards a modular, self-learning data quality assurance system.

Keywords:Anomaly detection, graph databases, Neo4j, automated data extraction

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back