Advances in artificial intelligence have enabled the widespread use of vector
embeddings to represent complex data, such as images, text and sound. As
modern applications generate large volumes of vector data, there is a need for
their efficient retrieval, storage and management. This problem is addressed
by vector databases. In this thesis, we analyse the impact of different index
structures and their configurations on search efficiency in vector databases.
Experiments were conducted in the Milvus vector database on the MS COCO
2017 image dataset, with vector embeddings generated using the CLIP model.
We compare the FLAT, HNSW and IVF_FLAT indices. We evaluate search
efficiency in terms of index construction time, retrieval time, query latency
and theoretical memory consumption. The results show that the choice of
appropriate index configuration parameters is crucial for search efficiency.
|