# How AI works > An interactive guide for engineers: how language models work under the hood and what matters when building agentic systems. With self-check questions. The guide is in English (Polish version under /pl/). Every topic: intuition, an interactive widget, mechanism, nuances, a short answer to check yourself, follow-up questions and sources. Last updated: 2026-09-26. What’s new: https://howaiworks.dev/changelog/ The whole guide in one file: https://howaiworks.dev/llms-full.txt Licence: text and images CC BY 4.0 (quote and adapt with a link to the source), code MIT. ## How a model reads and predicts From characters on screen to deciding what the next piece of text will be. - [Tokens](https://howaiworks.dev/tokens.md): A model sees no letters, only pieces of words turned into numbers. - [Vectors and matrices](https://howaiworks.dev/vectors-and-matrices.md): A token is a vector, and a layer multiplies it by fixed weight matrices. - [Attention](https://howaiworks.dev/attention.md): Every word looks back and chooses which earlier words to take information from. - [The next token](https://howaiworks.dev/sampling.md): The model gives odds for every token, and the sampler draws one. ## Where a model’s knowledge comes from Training, distillation, fine-tuning, the assistant role in a chat, and thinking before answering. - [Training](https://howaiworks.dev/training.md): Wide reading, vocational training, graded practice. - [Distillation](https://howaiworks.dev/distillation.md): How a small model learns from a large one: hard labels and distributions, on-policy, reasoning, limits and terms of service. - [Fine-tuning and LoRA](https://howaiworks.dev/fine-tuning.md): When to fine-tune and when a prompt or RAG is enough, how LoRA trains a fraction of the weights, and what fine-tuning breaks. - [How a model sees a chat](https://howaiworks.dev/chat-template.md): A conversation is one document, and the model writes the next line. - [Reasoning models](https://howaiworks.dev/reasoning-models.md): Scratch paper before the answer: it helps on multi-step tasks, and you pay for every line. ## How a model writes and what it costs Writing one piece at a time, caching, and why long conversations are expensive. - [The generation loop and KV cache](https://howaiworks.dev/kv-cache.md): The generation loop, prefill versus decode, and the GPU memory taken up by the key-value cache. - [Prompt caching](https://howaiworks.dev/prompt-caching.md): The KV cache kept by the provider between requests: cheaper input and a faster start, but only for an identical prompt prefix. - [The context window and agents](https://howaiworks.dev/context-window.md): What counts towards the limit, why each agent turn costs more than the last, and why quality drops with length. ## Context and knowledge What to put into the context window and how to find it in your data. - [Context engineering and memory](https://howaiworks.dev/context-engineering.md): What to put in the context window before each call, and how an agent remembers what doesn’t fit. - [Embeddings and vector search](https://howaiworks.dev/embeddings.md): Searching by meaning instead of by words: how it works, what it costs and where it fails. - [RAG](https://howaiworks.dev/rag.md): The pipeline from indexing and chunking, through retrieval and reranking, to citations, permissions and evaluation. ## Agents A model in a loop: tools, the standard for connecting them, enforced formats, the coding-agent harness and splitting the work. - [The agent loop](https://howaiworks.dev/agent-loop.md): A model in a loop picks the next step itself, until it decides it is done. - [Tools (function calling)](https://howaiworks.dev/tool-calling.md): The model doesn’t press buttons. It writes which one to press, and your code presses it. - [MCP](https://howaiworks.dev/mcp.md): One standard for connecting tools to many model-powered applications. The model still sees only definitions in the prompt. - [Enforcing output format](https://howaiworks.dev/structured-output.md): Constrained decoding: strict mode guarantees JSON that matches the schema, but not its content. - [The coding-agent harness](https://howaiworks.dev/coding-agent-harness.md): Everything around the model that turns it into a coding agent: tools, instructions, context, permissions and undo. - [Multiple agents](https://howaiworks.dev/multi-agent.md): A lead agent delegates parts of a task to subagents with clean contexts. Faster when the parts are independent, usually more expensive. ## Quality and security Hallucinations, measuring quality instead of guessing, and attacks through untrusted content. - [Hallucinations](https://howaiworks.dev/hallucinations.md): A model has no built-in “I don’t know”; it guesses in the same confident tone. - [Evals](https://howaiworks.dev/evals.md): A fixed set of cases run after every change, plus measurement in production, instead of “it seems better”. - [Prompt injection](https://howaiworks.dev/prompt-injection.md): Text the agent only reads can start steering it. Architecture defends against this, not the prompt. ## Production and serving A reliable system around the model, and what happens in the data centre. - [LLMs in production](https://howaiworks.dev/llms-in-production.md): What to build around a model API: retries, fallbacks, limits, costs, traces and safe rollouts. - [Why the GPU is idle](https://howaiworks.dev/roofline.md): During generation the GPU mostly waits for weights to arrive from memory. That explains batching, quantisation and pricier output tokens. - [Continuous batching](https://howaiworks.dev/continuous-batching.md): The server swaps conversations in and out of the batch at every step and gives them memory in blocks, so batch slots don’t sit empty. - [Speculative decoding](https://howaiworks.dev/speculative-decoding.md): A cheap draft guesses several tokens and the large model checks them in one pass. Faster, with the same output. - [Quantisation](https://howaiworks.dev/quantization.md): Weights in 8 or 4 bits instead of 16: 2–4 times less memory and faster generation for a small loss in quality. - [Mixture of Experts](https://howaiworks.dev/mixture-of-experts.md): Many experts, and a router picks a few per token: compute like a small model, memory like a large one. ## Choosing a model Matching a model to the task, open weights, alternatives to an LLM, and why a model sometimes seems “dumber”. - [Choosing a model](https://howaiworks.dev/choosing-a-model.md): The cheapest model that passes your eval set: cost per task, latency, limits and reading benchmarks. - [Open-weight models](https://howaiworks.dev/open-weight-models.md): Weights you can download and run yourself. Control over data and version, at the cost of GPUs and operations. - [When not to use an LLM: classifiers and System One](https://howaiworks.dev/when-not-to-use-an-llm.md): When a classifier, logprob scoring or a decision model beats an LLM. TypeSafe’s System One as the worked example, with vendor claims kept apart from facts. - [Is the model getting dumber?](https://howaiworks.dev/compute.md): Pinned weights do not change; the system around them does, and apps and aliases can switch to new weights. What has been documented, what is a hypothesis, and how to measure it. ## Practice and review Design scenarios and flashcards for review. - [Design a system](https://howaiworks.dev/system-design.md): Five real-world design scenarios: requirements, architecture, decisions, evaluation, failures. - [Flashcards](https://howaiworks.dev/flashcards.md): One question for every topic: answer it yourself first, then flip the card.