Details

Zasnova cloud native arhitekture za trajno hrambo vektoriziranih informacij agentnih sistemov
ID Šavron, Peter (Author), ID Jurič, Branko Matjaž (Mentor) More about this mentor... This link opens in a new window

.pdfPDF - Presentation file, Download (4,35 MB)
MD5: 83B4E3C794F8785D8BB744FCCF3CF396

Abstract
V delu obravnavamo problem učinkovite trajne hrambe in iskanja vloženih podatkov za potrebe agentov, ki temeljijo na velikih jezikovnih modelih. Ker neprestano večanje kontekstnega okna poslabša učinkovitost in poveča stroške, predlagamo oblačno arhitekturo, ki temelji na vektorski podatkovni bazi. Prek poznavanja osnovnih konceptov vektorskih podatkovnih baz ter analizo baz Milvus in Qdrant pridobimo znanje, ki nam omogoča zasnovati arhitekturo sistema, ki bo te podatke trajno hranil. Rešitev temelji na večnajemniški arhitekturi s Qdrantom, posrednikom Envoy, vmesnikom Quarkus in sistemom OpenFGA za zagotavljanje varnosti. Na podlagi obsežnega empiričnega testiranja različnih konfiguracij kazal, gruč in metod kvantizacije (kot je rotacijska kvantizacija) pokažemo, da lahko z optimizirano uporabo podgrafov najemnikov in ustreznih vrst kvantizacije bistveno zmanjšamo porabo virov ter izboljšamo prepustnost podatkovne baze, ne da bi pri tem žrtvovali priklic poizvedb.

Language:Slovenian
Keywords:vektorska podatkovna baza, vložitev, računalništvo v oblaku, kvantizacija, kazala, vmesniki, MCP, RAG, večnajemništvo, iskanje približnih najbližjih sosedov, merilo
Work type:Bachelor thesis/paper
Organization:FRI - Faculty of Computer and Information Science
Year:2026
PID:20.500.12556/RUL-187798 This link opens in a new window
Publication date in RUL:14.09.2026
Views:64
Downloads:16
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Design of a cloud native architecture for the persistent storage of vectorized information of agentic systems
Abstract:
In this work, we address the challenge of efficient persistent storage and retrieval of embedded data for large language model–based agents. Since continuously expanding the context window degrades performance and increases costs, we propose a cloud architecture built on a vector database. By examining the fundamental concepts of vector databases and analyzing Milvus and Qdrant, we establish the foundational knowledge needed to design the system architecture. Our solution implements a multi-tenant architecture that uses Qdrant, an Envoy proxy, a Quarkus interface, and OpenFGA for security. Through extensive empirical testing of various index configurations, cluster setups, and quantization methods (such as rotational quantization), our results show that optimizing the use of tenant subgraphs and selecting appropriate quantization techniques can significantly reduce resource consumption and improve database throughput without sacrificing query accuracy.

Keywords:vector database, embedding, cloud computing, quantization, index, interface, MCP, RAG, multi-tenancy, approximate nearest neighbor search, benchmark

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back