Glossary
AI glossary: LLM terms for developers
40 terms you meet when building with AI models, each defined in one plain sentence, then explained with a worked example, what it costs you, and what people confuse it with.
40 terms
Tokens and cost (8)
- API keysecret key · API tokenAn API key is a secret string that identifies your account to a service such as the OpenAI, Claude or Gemini API, so every request made with it is authorised, rate-limited and billed to you.
- Batch APIMessage Batches API · batch processing · batch modeA batch API is an asynchronous way to send many model requests as one job, which the provider works through when it has capacity, typically within 24 hours, at half the normal per-token price on OpenAI, Anthropic and Google.
- BYOKbring your own key · bring your own API keyBYOK (bring your own key) is a model where an AI app or tool runs on your own API key from a provider such as OpenAI, Anthropic or Google, so usage is billed to your account instead of the app’s.
- Prompt cachingcontext cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.
- Rate limitRPM and TPM limits · 429 Too Many RequestsA rate limit is a cap on how many requests or tokens an account may send to an API per minute or per day, and going over it makes the API reject requests with HTTP 429 until the allowance refills.
- Reasoning tokensthinking tokens · hidden reasoningReasoning tokens are the tokens a reasoning model generates while it works through a problem before writing its answer, billed as output tokens even though the API hides them or returns only a summary.
- TokenLLM token · AI tokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- Tokenizertokeniser · BPE tokenizerA tokenizer is the part of a language model that splits text into tokens from a fixed vocabulary, converts them to ID numbers for the model, and turns the model’s output IDs back into text.
Models and context (10)
- Chain of thoughtchain-of-thought prompting · CoT · step-by-step promptingChain-of-thought prompting is asking a language model to write out intermediate reasoning steps before its final answer, usually by showing worked examples or telling it to think step by step.
- Context windowcontext length · context size · token limitA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.
- Few-shot promptingmultishot prompting · in-context examples · k-shot promptingFew-shot prompting is giving a model a handful of worked input-and-output examples inside the prompt, so it copies the pattern for a new input without any retraining.
- HallucinationAI hallucination · LLM hallucinationA hallucination is a confident, plausible-sounding statement from a language model that is false or unsupported by its input, such as an invented citation, date, function or quote.
- Max output tokensmax_tokens · max_output_tokens · max_completion_tokens · maxOutputTokens · output token limitMax output tokens is the cap on how many tokens a model may generate in one response, set per request up to the model’s own limit, with any reasoning tokens counted inside it.
- Model snapshotpinned model version · dated model ID · model aliasA model snapshot is a fixed version of a model behind a specific API model ID, so requests to that ID keep getting the same model until it is retired, unlike an alias that the provider can move.
- System promptsystem message · system instructions · developer messageA system prompt is the set of standing instructions a developer sends with every request to set a model’s role, rules and output format, kept separate from what the user types.
- Temperaturesampling temperature · LLM temperatureTemperature is a sampling setting that divides a model’s next-token scores before they become probabilities, so low values make the likeliest token dominate and high values spread the choice across more tokens.
- Time to first tokenTTFT · first-token latencyTime to first token (TTFT) is how long a model takes from receiving a request to returning the first token of its reply, the wait a user feels before streamed text starts to appear.
- Top-ptop_p · nucleus sampling · topPTop-p, or nucleus sampling, is a sampling setting that keeps only the smallest set of likeliest next tokens whose probabilities add up to at least p, then picks the next token from that set.
Agents and tools (8)
- AI agentLLM agent · agentic loop · agent loopAn AI agent is a program in which a language model works towards a goal by repeatedly choosing a tool to call, reading the result and deciding the next step, until the task is done or a stop condition is reached.
- Claude Code hookshooks · PreToolUse hook · lifecycle hooksClaude Code hooks are commands, HTTP calls, MCP tool calls, prompts or subagents that Claude Code runs automatically at fixed points in a session, such as just before a tool call, so a rule is applied every time.
- CLAUDE.mdAGENTS.md · CLAUDE.local.md · project instructions fileCLAUDE.md is a Markdown file of project instructions that Claude Code loads into context at the start of every session, and AGENTS.md is the open, tool-neutral equivalent read by many coding agents.
- Function callingtool use · tool calling · tool callsFunction calling, also called tool use, is an LLM API feature in which the model replies with a structured request to run a function you described, with JSON arguments, which your code executes before sending the result back.
- MCPModel Context ProtocolMCP stands for Model Context Protocol, an open standard that lets AI applications such as Claude Code, ChatGPT or VS Code connect to outside tools and data through one common interface.
- OpenAI-compatible APIOpenAI compatibility · OpenAI-compatible endpoint · base_urlAn OpenAI-compatible API is a model API that accepts OpenAI’s Chat Completions request format, so you can call it with the official OpenAI SDK by changing only the base URL, the API key and the model name.
- Structured outputsstructured output · JSON Schema output · strict modeStructured outputs is an LLM API feature that constrains the model’s reply to a JSON Schema you supply, so the response parses and has the fields and types you asked for, unlike JSON mode, which only promises valid JSON.
- Subagentsub-agent · Claude Code subagent · custom agentA subagent is a separate AI worker that a main agent hands a task to, running in its own context window with its own system prompt, tools and model, and returning only its result to the main conversation.
Local AI (6)
- GGUFGGUF file · .gguf · GGUF formatGGUF is a single-file binary format from the ggml and llama.cpp project that stores a model’s weights together with everything needed to run them, such as its architecture, tokenizer and quantisation type.
- KV cachekey-value cache · attention cacheThe KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.
- Mixture of expertsMoE · sparse mixture of experts · MoE modelA mixture-of-experts (MoE) model splits parts of each layer into many parallel sub-networks called experts and uses a small router to send each token through only a few of them, so it computes with a fraction of its parameters.
- Open weightsopen-weight model · downloadable model weightsOpen weights means a model’s trained parameters are published for anyone to download and run on their own hardware, under a licence that sets what they may do with them.
- Quantisationquantization · LLM quantization · model quantisationQuantisation is storing a model’s weights in fewer bits than the 16 per weight most models are released with, so the model needs less memory and runs on smaller hardware, at a small cost in accuracy.
- VRAMvideo RAM · GPU memory · graphics memoryVRAM is the memory on a graphics card, and for running AI models locally it is the main limit: a model runs at full GPU speed only when its weights, KV cache and working buffers all fit in it.
Data and training (6)
- Chunkingtext chunking · document chunking · text splittingChunking is splitting documents into smaller pieces, usually a few hundred tokens each, before embedding them, so that a RAG system can search, retrieve and paste in only the parts relevant to a question.
- Embeddingsembedding · vector embeddings · text embeddingsEmbeddings are lists of numbers (vectors) that an embedding model produces from text, images or other data, placed so that inputs with similar meaning end up close together and can be compared mathematically.
- Fine-tuningfinetuning · supervised fine-tuning · SFT · model tuningFine-tuning is training an existing model further on a set of your own example inputs and ideal outputs, so its weights change and it follows a task, format or style without long instructions in every prompt.
- LoRAlow-rank adaptation · LoRA adapter · LoRA fine-tuningLoRA (low-rank adaptation) is a fine-tuning method that freezes a model’s original weights and trains two small matrices per adapted layer, whose product is added to the frozen weights, so only a tiny fraction of parameters is trained.
- RAGretrieval-augmented generation · retrieval augmented generationRAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.
- Vector databasevector store · vector DB · vector searchA vector database is a data store that keeps embeddings alongside their source data and quickly finds the stored vectors closest to a query vector, which is how semantic search and RAG fetch relevant text.
Web and SEO (2)
- AI crawlerAI bot · AI web crawler · AI user agent · LLM crawlerAn AI crawler is an automated bot that fetches web pages for an AI company, whether to collect training data, to index pages for AI search answers, or to read a page that a user asked an assistant about.
- llms.txtllms txt · LLMs.txt filellms.txt is a proposed convention for a Markdown file at a website’s /llms.txt path that gives language models a short summary of the site and curated links to its most useful pages.
Method
How these definitions are written
Each term opens with a single sentence you can quote on its own, then explains how the thing works, with an example that uses real numbers: token counts measured with the same tokenizers as the token counter, and prices taken from the daily data behind the cost calculator, so they stay current. Every page links the official documentation or paper it was checked against, and the date it was checked.
A term page defines; the guides teach. Where a topic needs a full walkthrough, such as prompt caching or running models locally, the term page links to the guide that covers it in depth.
Last checked 11 October 2026