<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=130550"><dc:title>Use of relaxed stochastic controls in reinforcement learning</dc:title><dc:creator>Rems,	Jan	(Avtor)
	</dc:creator><dc:creator>Agram,	Nacira	(Mentor)
	</dc:creator><dc:creator>Košir,	Tomaž	(Komentor)
	</dc:creator><dc:subject>reinforcement learning</dc:subject><dc:subject>exploration</dc:subject><dc:subject>stochastic control theory</dc:subject><dc:subject>relaxed controls</dc:subject><dc:subject>dynamical programming</dc:subject><dc:subject>optimal investment strategy</dc:subject><dc:description>In this work, we investigate how relaxed stochastic controls are used for exploration in continuous time and space reinforcement learning. The environment $X^u$ is modeled by a stochastic differential equation controlled by control $u$, while the value function $V^u$ is an infinite horizon performance functional. For relaxed control distribution $\pi$ we introduce relaxed versions of environment $X^{\pi}$ and value function $V^{\pi}.$ In a special linear-quadratic case the optimal control distribution turns out to be Gaussian with mean depending on the current state, and variance depending on exploration weight parameter. A reinforcement learning algorithm for optimal investment strategy in a simple model of the financial market with the infinite horizon is developed and tested.</dc:description><dc:date>2021</dc:date><dc:date>2021-09-16 08:15:15</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>130550</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
