AI glossary · Data and training
What are embeddings?
Also called: embedding, vector embeddings, text embeddings
Definition
Embeddings are lists of numbers (vectors) that an embedding model produces from text, images or other data, placed so that inputs with similar meaning end up close together and can be compared mathematically.
Explained
How it works
An embedding model reads an input and returns a vector of fixed length: OpenAI’s text-embedding-3-small returns 1,536 numbers, text-embedding-3-large returns 3,072, and Google’s gemini-embedding-2 returns 3,072 by default. No single number means anything on its own. What matters is position: texts about the same thing land near each other, even when they share no words.
To compare two embeddings you usually use cosine similarity: the dot product divided by the product of the two vectors’ lengths. Values near 1 mean very similar; values near 0 mean unrelated. OpenAI recommends cosine similarity and normalises its embeddings to length 1, so a plain dot product gives the same ranking. Embed your documents once, embed each question, return the closest matches: that is semantic search, and the retrieval half of RAG.
You can often request shorter vectors: OpenAI’s dimensions parameter and Gemini’s output_dimensionality (128 to 3,072, with 768, 1,536 or 3,072 recommended) shrink storage and speed up search. Only compare vectors from the same model at the same size, because different models place meaning differently.
Example
Cosine similarity with toy vectors
Real embeddings are too long to read, so here are made-up four-number vectors to show the arithmetic. Say “refund policy” embeds as [0.8, 0.1, 0.5, 0.2], “money back” as [0.7, 0.2, 0.6, 0.1] and “GPU drivers” as [0.1, 0.9, 0, 0.4]. The first pair scores 0.978 and the second 0.260, so a search for “money back” ranks the refund page first, even though the two phrases share no words.
A real model does the same sum over far more numbers. The question “Can I cancel my annual plan and get a refund for the unused months?” is 15 tokens in cl100k_base, the tokenizer OpenAI says to count with for its third-generation embedding models, and comes back as 1,536 numbers from text-embedding-3-small.
| Pair | Dot product | Cosine similarity |
|---|---|---|
| “refund policy” and “money back” | 0.90 | 0.978 |
| “refund policy” and “GPU drivers” | 0.25 | 0.260 |
Toy vectors, invented for illustration and not produced by any model. Cosine similarity computed at build time.
const dot = (a: number[], b: number[]) => a.reduce((s, x, i) => s + x * b[i], 0);
const cosine = (a: number[], b: number[]) =>
dot(a, b) / (Math.sqrt(dot(a, a)) * Math.sqrt(dot(b, b)));
cosine([0.8, 0.1, 0.5, 0.2], [0.7, 0.2, 0.6, 0.1]); // 0.978
cosine([0.8, 0.1, 0.5, 0.2], [0.1, 0.9, 0, 0.4]); // 0.260Cost and quality
Why it matters
Embeddings decide what a search or RAG system can find. If the right passage doesn’t sit near the question in vector space, no amount of prompt work afterwards can recover it, so the embedding model and how you chunk your text deserve as much testing as the chat model.
Embedding is billed per input token, separately from generation (our daily price data covers chat models, so check the provider’s pricing page). Size has a running cost too: every vector is stored and searched, and a 3,072-dimension vector takes twice the space of a 1,536-dimension one in a vector database.
Don’t mix up
Common confusions
- Embeddings vs tokens
- A token is a piece of text, and its ID is just a position in a vocabulary. The embeddings you store for search describe the meaning of a whole input (a sentence, a chunk, a page) and come from a separate embedding endpoint. Language models also have internal token embeddings, but those aren’t what a search index holds.
- Embedding model vs chat model
- An embedding endpoint returns vectors, never text. You can’t ask it a question; you use it to find what to show a chat model.
Go deeper
Try it and read more
- Free toolAI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Guide · 12 min readRAG chunk size and overlap: how to choose the right settingsHow to choose a RAG chunk size and overlap: token vs character splitters, embedding limits, what research shows, overlap cost maths and a test you can run.
Related
Related terms
- RAGRAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.
- Vector databaseA vector database is a data store that keeps embeddings alongside their source data and quickly finds the stored vectors closest to a query vector, which is how semantic search and RAG fetch relevant text.
- ChunkingChunking is splitting documents into smaller pieces, usually a few hundred tokens each, before embedding them, so that a RAG system can search, retrieve and paste in only the parts relevant to a question.
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- TokenizerA tokenizer is the part of a language model that splits text into tokens from a fixed vocabulary, converts them to ID numbers for the model, and turns the model’s output IDs back into text.