Details

Spodbujevano učenje na impulznih nevronskih mrežah
ID Pogačnik, Matjaž (Author), ID Bosnić, Zoran (Mentor) More about this mentor... This link opens in a new window

.pdfPDF - Presentation file, Download (1,46 MB)
MD5: B5EE8AC30AA4E26E61542926A0BD85AB

Abstract
V diplomskem delu obravnavamo spodbujevano učenje na impulznih nevronskih mrežah, ki se obnašajo podobno kot človeški možgani. Predstavimo in rešimo ključne izzive učenja impulznih mrež ter razvijemo rešitve, ki upoštevajo tako biološko smiselnost kot zahtevnost simulacije. S prilagoditvijo klasične oblike sinapse, odvisne od nagrajevanja in časovne razporeditve impulzov (R-STDP), omogočimo učinkovito učenje in dodeljevanje zaslug preteklim odločitvam, pri čemer ohranimo osnovni princip R-STDP, kjer sinapse kodirajo vpliv pretekle aktivnosti nevronov na izbrane akcije brez uvedbe negativnih nagrad ali nerealističnih mehanizmov. Razviti sistem razširimo v arhitekturo akter-kritik, ki omogoča reševanje problemov z zakasnjenimi nagradami s prenosom pričakovane nagrade iz cilja v prejšnja stanja. Sistem ovrednotimo na poenostavljenih in nato na kompleksnejših nalogah, kot sta igra Pong in problem mrežnega sveta (gridworld). V igri Pong rezultati kažejo postopno daljše sekvence igranja brez zgrešitve žogice, v nalogi mrežnega sveta pa postopno izboljševanje strategije in krajšanje poti do cilja.

Language:Slovenian
Keywords:impulzne nevronske mreže, spodbujevano učenje, R-STDP učenje, TD učenje
Work type:Bachelor thesis/paper
Typology:2.11 - Undergraduate Thesis
Organization:FRI - Faculty of Computer and Information Science
Year:2026
PID:20.500.12556/RUL-178161 This link opens in a new window
COBISS.SI-ID:268513283 This link opens in a new window
Publication date in RUL:20.01.2026
Views:305
Downloads:107
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Reinforcement learning in spiking neural networks
Abstract:
In this thesis, we study reinforcement learning in spiking neural networks, a type of neural network that behaves similarly to the human brain. We present and address key challenges in training spiking networks and develop solutions that account for both biological plausibility and simulation complexity. By adapting the classical form of reward modulated spike timing dependent plasticity (R-STDP), we enable efficient learning and the assignment of credit to past decisions, while preserving the core principle of R-STDP, in which synapses encode the influence of past neuronal activity on selected actions, without introducing negative rewards or unrealistic mechanisms. We then extend the developed system into an actor-critic architecture, enabling it to solve problems with delayed rewards by propagating the expected reward from the goal back to previous states. The system is first evaluated on simplified tasks and then on more complex problems, such as the game Pong and the gridworld problem. In Pong, the results show progressively longer sequences of play without missing the ball, while in the gridworld task the strategy gradually improves and the path to the goal becomes shorter.

Keywords:spiking neural networks, reinforcement learning, R-STDP learning, TD learning

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back