Guides
Guides for building with AI
Plain-English explanations of the things that decide what AI costs and how well it works: tokens, context, caching, models and agents. Every number is measured or sourced, and each guide links to a free tool so you can check it with your own data.
Tokens & costs
- What is a token in AI? A plain-English guide with real examplesWhat LLM tokens are, how text is split into them, why the same text costs different amounts on different models, and how to count tokens exactly.7 min ·
- How to estimate LLM API costs: the formula, worked examples and the trapsEstimate LLM API costs per request and per month: the token formula, worked examples for chatbots, RAG and agents, plus caching, batch and reasoning tokens.10 min ·
- The cheapest LLM APIs right now, ranked from daily price dataThe cheapest LLM APIs ranked by blended price from daily data, the cheapest for tools, vision and long context, free tiers, and how to cut cost per task.9 min ·
- Prompt caching explained: how it works on OpenAI, Anthropic and Gemini, and when it paysHow prompt caching works on OpenAI, Anthropic and Gemini: what gets cached, minimum lengths, how long caches last, write and read prices, and when it pays off.11 min ·
- Context windows explained: what counts, what happens at the limit, and how to check fitWhat an LLM context window is, what counts toward it, max output vs context, what happens when you exceed it, and when RAG beats long context.10 min ·
- How to count tokens in Python and JavaScript (OpenAI, Claude, Gemini, Qwen)Count LLM tokens in Python and JavaScript: tiktoken, gpt-tokenizer, Hugging Face tokenizers, and the official Claude, Gemini and OpenAI count endpoints.13 min ·
Models
Coding agents
- Claude Code pricing: plans vs API, and which costs lessWhich Claude plans include Claude Code, how the usage limits work, what an agent step costs on the API, and where Pro, Max and pay-as-you-go break even.11 min ·
- How to choose a model for coding agents: the criteria that matterPick an LLM for a coding agent on tool-use reliability, context, speed, price per task and where you can run it, with live prices for tool-capable models.10 min ·
- How to reduce Claude Code token usage, and what each fix costs youCut Claude Code token use with /clear, /compact, model and effort choice, a lean CLAUDE.md, fewer MCP servers and subagents, and the trade-off of each.11 min ·
- How to write a good CLAUDE.md, with a complete exampleWhat to put in CLAUDE.md, what to leave out, where Claude Code loads it from, how imports and AGENTS.md work, plus an annotated example and a before and after.12 min ·
- What is MCP (Model Context Protocol)? A plain-English guideWhat the Model Context Protocol is, how hosts, clients and servers talk, what tools, resources and prompts do, which apps support it and how to use it safely.11 min ·
- How to set up MCP servers in Claude CodeAdd MCP servers to Claude Code with claude mcp add, share them in .mcp.json, keep tokens out of git, pick the right scope and fix servers that won’t connect.12 min ·
Local AI
- How much VRAM do you need to run an LLM locally?VRAM needed for 8B to 70B models at Q4, Q8 and FP16, what fits on 8 to 80 GB GPUs, and how context length, KV cache and CPU offload change it.12 min ·
- How to run LLMs locally on a MacWhat fits on 16 to 128 GB Apple Silicon Macs, how much memory the GPU can use, how to set up Ollama, LM Studio, llama.cpp or MLX, and how fast it runs.12 min ·