IUT de Toulouse · Université de Toulouse · Tech Talk

Does AI lie?!

How – and how much – do vector databases improve AI performance?

Victor Simonet Unai Murillo

AI chatbots sometimes answer with total confidence… and get it completely wrong. In our group project, we looked at one popular solution: connecting the AI to a vector database so that it can look up real documents before answering. Below you will find our poster, followed by extra material from our tech talk.

See the poster ↓
Research poster 'Does AI lie?!' with the abstract, introduction, methods and a bar chart showing wrong answers before and after RAG for six AI models.
View full size Download PDF

Tip: on a phone, open the full-size version and pinch to zoom.

Going further

Beyond the poster

A poster has limited space. This part explains the ideas behind our research in more detail: what an embedding is, how a vector database finds information so quickly, and why retrieval is not a magic solution.

The poster in 30 seconds

The problem

LLMs hallucinate, their knowledge is frozen at training time, and retraining them is very expensive.

Our method

A literature review plus a benchmark: 6 AI models, 18 questions each, tested with and without retrieval.

Our answer

Vector databases cut wrong answers by 72 % on average – if the retrieval strategy fits the use case.

Key result

Giving the AI a document to read cuts its wrong answers by 72 %

Wrong answers, out of 100 – data from Sng, Zhang & Mueller (2024)

Before RAG After RAG
50 %average wrong answers without RAG
13.9 %average wrong answers with RAG
6 / 6models improved with retrieval

The weakest models gained the most: GPT-3.5 dropped from 72.2 % to 5.6 % wrong answers. Llama 3 improved the least and still made mistakes on more than a quarter of the questions, even with the right document in front of it.

What is an embedding?

Computers do not understand words, they understand numbers. An embedding is a list of numbers (a vector) that represents the meaning of a piece of text. It is produced by a neural network called an embedding model.

The key idea: texts with similar meanings get vectors that are close to each other. “Kitten” ends up near “cat”, far away from “truck”. Real embeddings have hundreds or thousands of dimensions; the drawing shows only two so that we can picture it.

To measure how close two vectors are, we usually compute the cosine similarity: the angle between them. A small angle means a similar meaning.

cat kitten dog car truck bus “pets?”
A question (◊) is turned into a vector too, so it lands next to the documents that talk about the same thing.

RAG, step by step

Retrieval-Augmented Generation happens in two phases.

Preparation (done once)

  1. Split the documents into small passages (“chunks”).
  2. Embed each chunk into a vector.
  3. Store the vectors in the vector database and build the index.

Answering (for every question)

  1. Embed the user's question.
  2. Retrieve the most similar chunks (the “top-k”).
  3. Augment the prompt: question + retrieved chunks.
  4. Generate: the LLM writes an answer based on these sources.

The big advantage: to update the AI's knowledge, you only add new documents to the database. No retraining is needed.

Popular vector databases

FAISS

Open-source library by Meta for fast similarity search, often used inside other systems.

Milvus

Open-source database designed for very large collections of vectors.

Qdrant

Open-source engine written in Rust, with filtering on metadata.

Weaviate

Open-source database that can combine vector and keyword search.

Chroma

Lightweight open-source option, popular for prototypes and small projects.

pgvector

Extension that adds vector search to the classic PostgreSQL database.

Pinecone

Fully managed cloud service: no server to install or maintain.

Limits: when RAG does not help

Our conclusion: vector databases are a cheap and effective way to make AI more reliable and up to date – provided the retrieval strategy is matched with the use case.

Glossary

LLM
Large Language Model: an AI trained on huge amounts of text to understand and generate language (GPT, Claude, Llama…).
Hallucination
A fluent and confident answer that is factually wrong.
Embedding
A vector of numbers that represents the meaning of a text.
Vector database
A database built to store embeddings and quickly find the most similar ones.
ANN
Approximate Nearest Neighbor: a fast search that finds almost the closest vectors.
RAG
Retrieval-Augmented Generation: the AI retrieves relevant documents before writing its answer.
Chunk
A small passage of a document, stored as one vector.
Recall
The share of relevant documents that the search actually finds.

References

Acknowledgments

We would like to thank our English teacher for their guidance throughout this project, and the IUT de Toulouse for giving us the opportunity to work on this topic.