This master's thesis examines large language models (LLMs) in business environments and proposes a knowledge management architecture that uses Retrieval-Augmented Generation (RAG) to limit unreliable answers and exploit unstructured internal data.
The system ingests data from document systems and communication platforms and transforms it into atomic facts and text chunks. It stores content and metadata in a relational database and their vector embeddings in separate vector-database collections. In adaptive retrieval, it first selects atomic facts or text chunks according to the type of question and, if necessary, retrieves additional context from the other collection.
We evaluated the pilot implementation on a synthetic corpus comprising 32 source documents organized into four chronological phases. The system divided the documents into 61 text chunks and extracted 420 atomic facts from them. The manual evaluation, blinded to strategy labels, covered 64 questions with one answer from each of four retrieval strategies and a no-RAG control, for 320 answers in total. Adaptive sequential retrieval, in which the agent chooses between facts and chunks, achieved the highest observed total score (5.83 out of 6.0) and, in one measurement run, the lowest median request elapsed time among the RAG strategies. The comparison did not fully separate the effect of retrieval strategy from differences in prompt wording, and the primary tests found no statistically significant differences between the four RAG strategies. A review of 50 active facts found high proportions of top ratings for atomicity, self-containment, and source fidelity, while all three supersessions in the final collection were substantively justified.
The results support technical feasibility in the synthetic pilot and indicate a possible trade-off among quality, elapsed time, and token usage, but broader repeated tests are needed before claiming superiority, scalability, security, or production readiness.
|