Going further
Beyond the poster
A poster has limited space. This part explains the ideas behind our research in more detail: what an embedding is, how a vector database finds information so quickly, and why retrieval is not a magic solution.
The poster in 30 seconds
LLMs hallucinate, their knowledge is frozen at training time, and retraining them is very expensive.
A literature review plus a benchmark: 6 AI models, 18 questions each, tested with and without retrieval.
Vector databases cut wrong answers by 72 % on average – if the retrieval strategy fits the use case.
Key result
Giving the AI a document to read cuts its wrong answers by 72 %
Wrong answers, out of 100 – data from Sng, Zhang & Mueller (2024)
The weakest models gained the most: GPT-3.5 dropped from 72.2 % to 5.6 % wrong answers. Llama 3 improved the least and still made mistakes on more than a quarter of the questions, even with the right document in front of it.
What is an embedding?
Computers do not understand words, they understand numbers. An embedding is a list of numbers (a vector) that represents the meaning of a piece of text. It is produced by a neural network called an embedding model.
The key idea: texts with similar meanings get vectors that are close to each other. “Kitten” ends up near “cat”, far away from “truck”. Real embeddings have hundreds or thousands of dimensions; the drawing shows only two so that we can picture it.
To measure how close two vectors are, we usually compute the cosine similarity: the angle between them. A small angle means a similar meaning.
How does a vector database search so fast?
The simple method is to compare the question with every vector in the database. This is exact, but far too slow when there are millions of documents. Vector databases therefore use Approximate Nearest Neighbor (ANN) search: they accept a tiny loss of precision in exchange for a huge gain in speed.
HNSW
Hierarchical Navigable Small World
Vectors are linked in a graph with several layers, a bit like motorways and small roads. The search starts on the top layer with long jumps, then zooms in on the lower layers.
IVF
Inverted File Index
Vectors are sorted into clusters. At search time, the database only looks inside the few clusters closest to the question instead of the whole collection.
PQ
Product Quantization
Vectors are compressed into short codes. They take much less memory, so billions of them can fit on a single machine, at the cost of a little accuracy.
The main trade-off is always the same: speed and memory versus recall (the share of truly relevant documents that are actually found).
RAG, step by step
Retrieval-Augmented Generation happens in two phases.
Preparation (done once)
- Split the documents into small passages (“chunks”).
- Embed each chunk into a vector.
- Store the vectors in the vector database and build the index.
Answering (for every question)
- Embed the user's question.
- Retrieve the most similar chunks (the “top-k”).
- Augment the prompt: question + retrieved chunks.
- Generate: the LLM writes an answer based on these sources.
The big advantage: to update the AI's knowledge, you only add new documents to the database. No retraining is needed.
Popular vector databases
FAISS
Open-source library by Meta for fast similarity search, often used inside other systems.
Milvus
Open-source database designed for very large collections of vectors.
Qdrant
Open-source engine written in Rust, with filtering on metadata.
Weaviate
Open-source database that can combine vector and keyword search.
Chroma
Lightweight open-source option, popular for prototypes and small projects.
pgvector
Extension that adds vector search to the classic PostgreSQL database.
Pinecone
Fully managed cloud service: no server to install or maintain.
Limits: when RAG does not help
- Best-case benchmark. In the study, the models were given the correct document. In real life, the search may return a wrong or incomplete passage, and the AI may then repeat that mistake with confidence.
- Small sample. 18 questions per model is enough to show a trend, but not to rank the models precisely.
- Too much context. Adding many documents is not always better: models tend to use information at the beginning and end of a long prompt better than information in the middle (“lost in the middle”).
- Strong models gain less. Models that already know the answer have less room for improvement, and for some tasks irrelevant context can even make them worse.
- Quality in, quality out. The answer can only be as good as the documents in the database: outdated or biased sources give outdated or biased answers.
Our conclusion: vector databases are a cheap and effective way to make AI more reliable and up to date – provided the retrieval strategy is matched with the use case.
Glossary
- LLM
- Large Language Model: an AI trained on huge amounts of text to understand and generate language (GPT, Claude, Llama…).
- Hallucination
- A fluent and confident answer that is factually wrong.
- Embedding
- A vector of numbers that represents the meaning of a text.
- Vector database
- A database built to store embeddings and quickly find the most similar ones.
- ANN
- Approximate Nearest Neighbor: a fast search that finds almost the closest vectors.
- RAG
- Retrieval-Augmented Generation: the AI retrieves relevant documents before writing its answer.
- Chunk
- A small passage of a document, stored as one vector.
- Recall
- The share of relevant documents that the search actually finds.
References
- Sng, Zhang & Mueller (2024). Benchmark of LLM accuracy with and without Retrieval-Augmented Generation.
- Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020.
- Malkov, Y. & Yashunin, D. (2020). Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE TPAMI.
- Johnson, J., Douze, M. & Jégou, H. (2019). Billion-Scale Similarity Search with GPUs. IEEE Transactions on Big Data.
- Liu, N. F. et al. (2024). Lost in the Middle: How Language Models Use Long Contexts. TACL.
Acknowledgments
We would like to thank our English teacher for their guidance throughout this project, and the IUT de Toulouse for giving us the opportunity to work on this topic.