AI glossary · Data and training
What is a vector database?
Also called: vector store, vector DB, vector search
Definition
A vector database is a data store that keeps embeddings alongside their source data and quickly finds the stored vectors closest to a query vector, which is how semantic search and RAG fetch relevant text.
Explained
How it works
Comparing a question’s embedding with every stored vector (exact, or brute-force, search) is simple and works for small collections, but the work grows with every vector you add. Vector databases add an approximate nearest neighbour (ANN) index that checks only a small, promising part of the collection, trading a little recall for much faster search.
A common index is HNSW (Hierarchical Navigable Small World), from a 2016 paper (arXiv 1603.09320): a layered graph that search enters at a sparse top layer and descends, which the authors say gives logarithmic complexity scaling. Stores also keep metadata such as source URL, date and permissions, so you can filter results and cite sources.
It doesn’t have to be a separate product. pgvector, an open-source PostgreSQL extension, adds a vector type with cosine (<=>), Euclidean (<->) and negative inner-product (<#>) distance, plus HNSW and IVFFlat indexes. Without an index it runs exact search.
Example
Storing 1,000,000 embeddings
OpenAI’s text-embedding-3-small returns 1,536 numbers per chunk. pgvector stores each vector in 4 bytes per dimension plus 8, so 6,152 bytes, and 1,000,000 chunks need 6.15 GB for the vectors alone, before the index, the chunk text and metadata.
The halfvec type stores 2 bytes per dimension and roughly halves that. Bigger models change the maths: text-embedding-3-large returns 3,072 dimensions, more than the 2,000 pgvector can index as vector, so you would index it as halfvec (up to 4,000) or request fewer dimensions.
| Dimensions and type | Bytes per vector | 1,000,000 vectors | HNSW index? |
|---|---|---|---|
1,536, vector | 6,152 | 6.15 GB | Yes |
1,536, halfvec | 3,080 | 3.08 GB | Yes |
3,072, vector | 12,296 | 12.30 GB | No: too many dimensions |
3,072, halfvec | 6,152 | 6.15 GB | Yes |
Storage formulas and index limits from the pgvector README, checked 2026-10-11. GB = 10⁹ bytes. Index, text and metadata are extra.
Cost and quality
Why it matters
Storage and memory grow with chunks times dimensions, so chunking and embedding size settle most of the bill before you choose a product. Approximate indexes can miss a true nearest neighbour, so measure recall on your own questions after tuning an index for speed.
Choose on scale, filtering needs and what you already run. If your data lives in PostgreSQL, an extension may be enough; a dedicated service adds managed scaling. Either way, the vectors only find what the embedding model put near the question, so retrieval quality is tested in the RAG pipeline as a whole.
Don’t mix up
Common confusions
- Vector database vs embedding model
- The database stores and searches vectors; the embedding model creates them. Some databases can call a model for you, but the vectors still come from an embedding model, and changing that model means re-embedding everything.
- Approximate vs exact search
- An ANN index returns the probable nearest neighbours, quickly. Exact search checks every vector and always returns the true nearest ones. pgvector’s README says adding an index “trades some recall for speed”.
Go deeper
Try it and read more
- Free toolAI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Guide · 12 min readRAG chunk size and overlap: how to choose the right settingsHow to choose a RAG chunk size and overlap: token vs character splitters, embedding limits, what research shows, overlap cost maths and a test you can run.
Related
Related terms
- EmbeddingsEmbeddings are lists of numbers (vectors) that an embedding model produces from text, images or other data, placed so that inputs with similar meaning end up close together and can be compared mathematically.
- RAGRAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.
- ChunkingChunking is splitting documents into smaller pieces, usually a few hundred tokens each, before embedding them, so that a RAG system can search, retrieve and paste in only the parts relevant to a question.