Modern AI Engineering

Vector Databases and Embeddings Explained for Beginners

By Modern AI Engineering · · 7 min read

An embedding is a list of numbers that represents the meaning of a piece of text, an image or another item. A vector database stores many embeddings and can quickly find the ones closest to a given embedding. Together they let software search by meaning instead of by exact words.

This is the technology under semantic search, recommendations and RAG. The article explains both halves and how they connect, with no heavy math.

What is an embedding?

Computers compare numbers easily and meanings poorly. An embedding closes that gap. An embedding model reads a piece of text and returns a vector, a long list of numbers, often hundreds or thousands of them.

The numbers are not readable one by one. What matters is position. You can think of each vector as a point in a space with many dimensions. The model was trained so that texts with similar meaning land near each other. A sentence about a cat sleeping on a sofa lands close to a sentence about a kitten resting on a couch, though they share almost no words.

The same idea works for images, audio and products. Anything a model can turn into a vector can be compared this way.

How do you measure similarity between vectors?

Once items are points in space, similar means close. There are three common ways to measure closeness.

MeasureWhat it comparesTypical use
Cosine similarityThe angle between two vectors, ignoring their length.Text embeddings. The usual default.
Dot productAngle and length together.Models trained with it. Equal to cosine when vectors have length one.
Euclidean distanceThe straight-line distance between two points.Cases where the size of the values carries meaning.

What is a vector database?

A vector database is a store built for one main question: which stored vectors are closest to this one? You give it a query vector and a number, say five, and it returns the five nearest items.

Each record usually holds three things: the vector, the original text or a pointer to it, and metadata such as the source, the date or the author. Metadata lets you filter. You can ask for the nearest chunks that also come from one product manual or from the current year.

Some vector databases are separate products. Others are an extension to a database you may already use. For a small project, a library that keeps vectors in memory is enough. The idea is the same in each case.

How does vector search work at scale?

The simplest search compares the query with every stored vector and keeps the best. This is exact, and it is fine for a small collection. With millions of vectors it becomes too slow for a live product.

Large systems use approximate nearest neighbour search, or ANN. It gives up a little accuracy to gain a great deal of speed. Instead of checking everything, it uses an index, a structure built ahead of time that leads the search towards the right region.

A widely used index is HNSW. It links each vector to some of its neighbours, forming a graph with several layers. The top layer has few points and long links, for big jumps. Lower layers have more points and shorter links, for fine steps. A search starts at the top and moves closer at every layer. Another approach groups vectors into clusters and searches only the nearest clusters.

Every index has settings that trade speed and memory against recall, which is the share of the true nearest items that the search finds.

Vector database vs traditional database

The two answer different kinds of question, and most products need both.

Traditional databaseVector database
Typical questionWhich rows match this exact value?Which items are most similar to this one?
DataTables of numbers, text and datesVectors with attached metadata
Match typeExactNearest, ranked by similarity
IndexSorted trees and hash tablesGraph or cluster indexes for ANN
ResultAll rows that matchThe top few closest items

How do embeddings and vector databases power RAG?

RAG is the most common reason people meet this topic. The flow is short.

  1. Split your documents into chunks.
  2. Send each chunk through an embedding model and store the vector with the text.
  3. When a user asks a question, embed the question with the same model.
  4. Ask the vector database for the nearest chunks.
  5. Put those chunks into the prompt, so the language model can answer from them.

Common mistakes with embeddings and vector search

These problems account for many poor search results.

  • Using one embedding model for the documents and a different one for the questions. Vectors from different models cannot be compared.
  • Changing the embedding model without embedding every document again.
  • Embedding chunks that are too large, so each vector is a blur of several topics.
  • Relying on vector search alone for exact terms such as codes and names. Add keyword search.
  • Embedding the same text again and again. A cache of embeddings saves time and money.
  • Tuning the index for speed without checking how much recall was lost.
  • Assuming that near in vector space means correct. Similar text can still be the wrong answer.

Do you always need a vector database?

No. If you have a few hundred or a few thousand chunks, you can keep the vectors in memory and compare the query with all of them. That is exact and fast enough. A dedicated vector database earns its place when the collection is large, when many users search at once, or when you need filtering, updates and backups handled for you.

It is also worth asking whether vectors are the right tool. Keyword search is hard to beat for exact terms, and some retrieval methods use no embeddings at all.

In the AI Engineering Bootcamp, the lesson on embeddings sits in Module 4, and Module 10 covers vector databases, ANN search, semantic search, hybrid search and embedding caches. A later lesson shows how images become embeddings too.

Frequently asked questions

What is a vector database in simple terms?

It is a database that stores lists of numbers called vectors and finds the ones closest to a query vector. Because close vectors mean similar content, it lets you search by meaning.

What is the difference between an embedding and a vector?

A vector is any list of numbers. An embedding is a vector produced by a model so that its position reflects the meaning of the item it represents.

Is a vector database required for RAG?

No. Small collections can be searched in memory, and some RAG systems use keyword search or other methods. A vector database helps when the collection or the traffic is large.

What is cosine similarity?

It is a measure of how closely two vectors point in the same direction. A higher value means the two items are more alike in meaning.

What does HNSW mean?

Hierarchical Navigable Small World. It is a graph index that lets a search jump quickly towards the nearest vectors without checking every one.

Learn it properly: the AI Engineering Bootcamp

More articles