<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=151240"><dc:title>Fashion Image Editing Through Text Descriptions</dc:title><dc:creator>Stopar,	Julija	(Avtor)
	</dc:creator><dc:creator>Štruc,	Vitomir	(Mentor)
	</dc:creator><dc:creator>Omachi,	Shinichiro	(Komentor)
	</dc:creator><dc:subject>diffusion models</dc:subject><dc:subject>text-to-image generation</dc:subject><dc:subject>human body segmentation</dc:subject><dc:subject>inpainting</dc:subject><dc:subject>virtual try-on</dc:subject><dc:description>As the performance of text-to-image generative neural networks keeps improving, much attention has been paid to potential uses of this technology. One industry that currently has great interest in innovations through artificial intelligence is the fashion industry. However, applications with more practical useability and an ability to provide assistance to creatives rather than threatening to replace them need to be designed carefully and include imposing additional constraints to out-of-the-box generative models. In addition to that, special care must be taken when designing applications that include depictions of people, as this raises a number of ethical as well as purely technical and aesthetic concerns due to the complexity of this task. For these reasons, we choose to instead combine the capabilities of Stable Diffusion, a diffusion probabilistic model capable of generating highly convincing images based on textual descriptions, with the principle of virtual try-on, a practice popularized in e-commerce which aims to realistically edit input images of people by changing the clothes they are wearing while preserving the rest of the image, most importantly the wearer. This thesis presents a possible implementation of a text-to-image generating pipeline, which provides the user with a photorealistic depiction of clothing, described in text, worn by the model whose image was provided as an input into the system, by editing only certain regions of an image using an inpainting technique. We design a robust framework, where in-the-wild images are also supported, capable of generating a wide range of clothing types with varying styles and silhouettes, which allows for creative use by designers and potential fashion customers alike. A key contribution of our approach is also that the end user does not have to provide additional input data (such as a mask, pose information, etc.) aside from the input image and textual description, written in natural language, as the area to be modified is determined automatically through human body segmentation using the DensePose algorithm. We display and comment on a large number of successful and less successful instances of produced images, identify the strengths and limitations of our approach, present the results of an anonymous survey that aims to evaluate the generated images using public perception, and compare our results to those obtained by a similar application developed prior to ours both qualitatively and quantitatively using a CLIP score. Although the findings of our experimentation are generally encouraging and the performance of the model is fairly consistent in terms of text-image alignment and the perceived realism of the generated images, we finally reflect on possible improvements to the application based on the common errors discovered through observing the results.</dc:description><dc:date>2023</dc:date><dc:date>2023-10-02 08:20:00</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>151240</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
