<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=171822"><dc:title>Detection of Face Forgeries with Self-Supervised Learning</dc:title><dc:creator>Todorov,	Leon	(Avtor)
	</dc:creator><dc:creator>Peer,	Peter	(Mentor)
	</dc:creator><dc:creator>IVANOVSKA PRESKAR,	MARIJA	(Komentor)
	</dc:creator><dc:subject>computer vision</dc:subject><dc:subject>attack detection</dc:subject><dc:subject>face image morphing attacks</dc:subject><dc:subject>face forg- eries</dc:subject><dc:subject>deep learning</dc:subject><dc:subject>self-supervised learning</dc:subject><dc:description>In biometric identity verification, the growing fidelity of AI-synthesized
faces is threatening the reliability of face recognition systems. Modern gen-
erative models can create extremely realistic facial forgeries that often evade
detection by current methods. Among these, morphing attacks are especially
dangerous: by digitally merging the faces of two or more individuals, they
yield a single image that can fool face recognition systems into matching
multiple identities, thereby facilitating identity fraud and other malicious
exploits. Most Morphing Attack Detection (MAD) approaches use super-
vised learning on a fixed set of known morphing techniques. Such models
often achieve high accuracy on morphs created by the same algorithms seen
during training, but they tend to rely on method-specific artifacts and strug-
gle to generalize to morphs from unseen techniques or under different data
conditions. Unsupervised one-class methods avoid overfitting to specific at-
tacks, but they often lack the sensitivity to detect the faint, distributed arti-
facts left by high-quality morphs. To overcome these challenges, a new face
forgery detection framework with improved robustness and generalization is
introduced. The approach uses self-supervised training on synthetic forgery
artifacts, which helps the detector learn decision boundaries that are more
generic and resilient. At the heart of the model, a gating mechanism fuses
two complementary information streams: the high-level semantic features
from a vision-language foundation model and the fine-grained spatial features
from a high-resolution convolutional network. The semantic branch is built
on a CLIP vision-language backbone and fine-tuned via Low-Rank Adap-
tation (LoRA) to adapt its image-text embeddings to the forgery detection
task, enabling accurate discrimination between authentic and manipulated
faces. In parallel, a high-resolution convolutional branch (based on HRNet)
preserves detailed spatial information and aggregates multi-scale features,
allowing it to capture even very subtle artifacts. An auxiliary segmentation
module provides pixel-level guidance to this branch by distinguishing genuine
facial regions from likely manipulated ones, which regularizes the training.
By combining the CLIP branch’s global semantic context with the convolu-
tional branch’s local artifact sensitivity, the model produces a well-balanced
and highly discriminative representation for detecting face forgeries. The
entire architecture is trained end-to-end with a composite loss that simulta-
neously enforces semantic alignment, segmentation consistency, and classifi-
cation accuracy. Evaluated on diverse morphing benchmarks, the proposed
method achieves state-of-the-art performance. It significantly outperforms
both supervised and unsupervised baseline detectors, attaining an average
Equal Error Rate (EER) of just 0.85%. Notably, the improvements are most
pronounced on high-quality morphs generated by advanced GAN and diffu-
sion models, highlighting the framework’s resilience against next-generation
forgery techniques.</dc:description><dc:date>2025</dc:date><dc:date>2025-09-03 08:25:01</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>171822</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
