Details

Semantično vodeni difuzijski modeli za neskončno velike toroidne teksture
ID KREFT, JAKOB (Author), ID Perš, Janez (Mentor) More about this mentor... This link opens in a new window, ID Lesar, Žiga (Comentor)

.pdfPDF - Presentation file, Download (40,18 MB)
MD5: D1AC56D6835F8AB543B6C8739F09E7A8

Abstract
Magistrsko delo obravnava problem generiranja velikih in brezšivnih tekstur lubja smrekovih hlodov. Generiranje takih tekstur otežujeta dve dejstvi. Prvič, hlod je po obodu valjast in v visoki ločljivosti lahko slikamo le ozek pas. Drugič, difuzijske modele, ki so danes prvi izbor za sintezo slik, učimo na zaplatah velikosti 512 x 512 pikslov, medtem ko obod hloda meri več tisoč pikslov. Cilj naloge je doučiti generativni model, ki iz omejenih posnetkov sestavi teksturo po celotnem obodu, jo brezšivno sklene v torus in pri tem ohrani naravno razporeditev razredov: lubje, slepice in mehanske poškodbe. Predlagamo dvofazni postopek. V prvi fazi z markovskimi naključnimi polji generiramo grob semantični zemljevid, ki nadzira globalne deleže razredov. V drugi fazi obstoječi latentni difuzijski model po zgledu pristopa DiffInfinite prilagodimo in doučimo na lastni učni množici, da iz semantičnega zemljevida ustvari sliko visoke ločljivosti. Pristop razširimo z modularnim indeksiranjem zaplat in periodičnim mešanjem pri dekodiranju, da je generirana tekstura brezšivno ploščitljiva po izbrani geometriji (toroidni ali valjni). Arhitektura in uteži difuzijskega modela pri tem ostanejo nespremenjene. V sodelovanju z Biotehniško fakulteto smo na terenu posneli 140 fotografij 15 smrekovih hlodov in iz njih pripravili podatkovni nabor 270 parov slik lubja in semantičnih mask. Difuzijski model smo doučili v 200,000 korakih; z razširjenim bogatenjem učnih zaplat smo dosegli občutno nižji vrednosti metrik FID in KID, ki smo ju v zaključni fazi z blažjim bogatenjem še dodatno znižali. Med učenjem smo ugotovili, da kriterijska funkcija po nekaj tisoč korakih ne odraža več napredka v vizualni kakovosti; metriki FID in KID, izračunani na zaplatah, pa nadaljnji napredek dobro ujameta. Generirano teksturo smo brezšivno ovili okoli demonstracijskega valjastega hloda v okolju Blender.

Language:Slovenian
Keywords:strojno učenje, difuzijski modeli, generativni modeli, teksture lubja, lubje hlodov, toroidne teksture, panoramske slike, neskončna ločljivost, semantična segmentacija
Work type:Master's thesis/paper
Organization:FE - Faculty of Electrical Engineering
Year:2026
PID:20.500.12556/RUL-185875 This link opens in a new window
Publication date in RUL:21.08.2026
Views:27
Downloads:15
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Semantically Guided Diffusion Models for Infinitely Large Toroidal Textures
Abstract:
This thesis addresses the problem of generating large, seamless bark textures of spruce logs. The problem arises from two facts. First, a log is cylindrical along its circumference, and only a narrow strip can be photographed at high resolution. Second, diffusion models, the current state of the art for image synthesis, are typically trained on patches of 512 x 512 pixels, while the full circumference spans several thousand. The goal is to fine-tune a generative model that composes the full circumferential texture from limited captures, closes it seamlessly into a torus, and preserves a natural distribution of classes: bark, knots, and mechanical damage. We propose a two-stage pipeline. In the first stage, a Markov Random Field generates a coarse semantic map that controls the global class proportions. In the second stage, an existing latent diffusion model based on the DiffInfinite approach is adapted and fine-tuned on our data to produce a high-resolution image conditioned on that map. We extend the approach with modular patch indexing and periodic blending during decoding, so that the generated texture tiles seamlessly under the chosen geometry, either toroidal or cylindrical. The U-Net architecture itself is unchanged. Together with the Biotechnical Faculty we captured 140 photographs of 15 spruce logs in the field and built a dataset of 270 bark image and mask pairs. We fine-tuned the diffusion model for 200.000 steps; broad augmentation of the training patches lowered FID and KID substantially, and lighter augmentation in the final phase lowered them further. We also observed that after a few thousand steps the loss function no longer reflects visual quality progress, while FID and KID computed on patches continue to track it well. The generated texture was wrapped seamlessly around a demonstration cylindrical log in Blender.

Keywords:machine learning, diffusion models, generative models, bark textures, log bark, toroidal textures, panoramic images, infinite resolution, semantic segmentation

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back