What is context engineering?
Context engineering is the practice of deciding what text goes into a large language model (LLM) on each call, and what stays out. That text fills the model's context window, the amount of text, measured in tokens, that the model can take into account at once. In a coding agent, the window holds instructions, files, tool results, and history.
Anthropic's engineering guide defines context engineering as "the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference." The guide states the aim as finding "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome." Anthropic's engineers describe context engineering as the natural progression of prompt engineering. An agent's context changes on every turn of its loop, not only at the first prompt.
Project rules and earlier decisions reach a model only through its context. When the context misses a rule, an agent can break the rule and still report the task as "done."
How does context engineering work?
Before each model call, the agent harness builds the context window, usually from these parts:
- System prompt. The harness's own instructions set how the model responds and uses tools.
- Tool definitions. Each tool has a name, a description, and an argument schema. Some harnesses load only the names of Model Context Protocol (MCP) tools until a task needs one, which saves room in the window.
- Instruction files. A file in the repository, e.g. AGENTS.md, gives the project's commands and rules.
- The task. The developer's request can point to a spec or an issue.
- Files and tool results. Each file read and each command's output arrives as the result of a tool call.
- History. Earlier messages and results stay until the harness removes or summarizes them.
The window grows on each turn as the harness adds each tool call and its output to the history. Context engineering sets what enters the window and when it leaves. The main techniques are these:
- Load at the start. Short, stable facts go in before the first call, e.g. the test command.
- Retrieve on demand. The agent keeps references, e.g. file paths, and reads a file only when the task needs it. Retrieval-augmented generation (RAG), a related technique, adds passages from a searched document index to the prompt.
- Compact. The harness replaces older history with a summary as the window nears its limit.
- Isolate. A subagent does one part of the task in its own context window and returns only a summary to its parent.
- Write notes. The agent saves notes to a file and reads them back after the history is gone.
Content that the agent reads, e.g. a fetched web page, can hide an attacker's instructions, which is called prompt injection.
What is an example of context engineering?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme asks a coding agent to "Let customers edit their delivery address during checkout." The first run uses the context that the harness loads by default:
- The harness loads its system prompt, the tool definitions, the project's
AGENTS.md, and the request. - The agent searches the codebase for
deliveryAddressand reads the 3 files that match. The cart code incart/store.jsis not among them. - The agent edits the address handler and runs
npm test, and 42 tests pass. - The agent's final message says the task is "done."
The developer tries the checkout. The address saved, but the cart emptied. Nothing in the agent's context said that an address change must keep the cart, and no test checked it.
The developer changes the context instead of rewording the request. Two lines go into the project's AGENTS.md:
Checkout: changing the delivery address must not change the cart.
Cart state lives in cart/store.js. Read it before you edit checkout code.
On the next run, the rule and the file path are in the context from the first call. By default, Claude Code reads AGENTS.md only when the project has no CLAUDE.md, or when CLAUDE.md imports it with @AGENTS.md. The agent reads cart/store.js before it edits the handler, and the handler it writes keeps the cart's items.
This example is simplified. A real project would also add a test that checks the cart after an address change, because an instruction in the context is not a check.
What happens when context runs out?
A long session fills the window with file reads, test output, and old messages. The /context command in Claude Code shows what fills it. The harness or the developer has to make room, and each way loses detail:
- Compaction. Claude Code's documentation says Claude Code compacts automatically as the window nears its limit, and its
/compactcommand does the same on request. - Trimming. The harness drops old tool results or the oldest messages.
- A fresh session. The developer starts again with a fresh window, e.g. with the
/clearcommand in Claude Code.
After compaction, Claude Code reads the project's root CLAUDE.md file from disk again. A rule typed once in chat is summarized with everything else, so it can drop out and the agent can break it later.
Quality can also drop before the window is full. Anthropic's guide describes an effect called "context rot." As the number of tokens in the window grows, the model's ability to recall information from it decreases. Claude Code's memory documentation also says that longer instruction files reduce adherence.
How does intent get lost between sessions?
Intent is the reason behind a change. When a coding agent writes the change, much of that reason is stated in the conversation that produced it. Each session starts with a fresh context window, which holds only what the harness loads at the start.
Unless someone wrote the reasons down, the next session gets the code but not the reasons for it. The diff shows that a workaround exists, not that a developer asked for it, so a later agent can remove it as unneeded code.
Teams carry intent forward in files that a later session can load:
- Instruction files. A rule that a developer types into chat twice belongs in the project's instruction file, which loads at the start of each session.
- Specs. Under spec-driven development, a written spec states what a feature should do, and the agent reads it before it plans.
- Saved sessions. AI coding session history keeps past prompts and replies, so a person or an agent can look up why a change was made.
- Bug reports. A bug report for an agent states the reproduction steps and the expected result, because the session that wrote the code is over.
How is context engineering different from prompt engineering?
Prompt engineering is the practice of wording a prompt, usually a system prompt or a request, so the model responds well. Context engineering covers the whole window on every call, and the prompt is one part of it. The two overlap, and a clear prompt still fails when the file that the task needs never reaches the model.
The two differ on these points:
| Aspect | Prompt engineering | Context engineering |
|---|---|---|
| What it shapes | The wording of one prompt | Everything in the window |
| When it happens | Before the first call | Before each call in the loop |
| Main material | Instructions and examples | Instructions, tools, files, results, and history |
| Typical failure | A vague request | A missing file, a lost rule, or a crowded window |
How does SpecStory help with context engineering?
The reasons behind earlier changes reach a later session only when they are saved where a person or an agent can read them. SpecStory keeps the conversations behind your code. Your sessions become local Markdown files you can read, search, and keep alongside your code, so a later agent session can read the prompts and decisions too. Saving a conversation with SpecStory does not itself test the software.
FAQs
Does a bigger context window remove the need for context engineering?
A bigger context window does not remove the need for context engineering. The model's recall from a long context gets weaker as the context grows, and a long session still fills the window. The agent also needs the right files and rules in the window, which size alone does not supply.
Is context engineering the same as retrieval-augmented generation?
Context engineering is broader than retrieval-augmented generation (RAG). RAG is one technique that searches an index of documents and adds matching passages to the prompt. Context engineering also covers instructions, tool definitions, history, and compaction, and coding agents often read files on demand instead.
How do subagents help with context?
Subagents help with context by doing one part of a task in their own context window. A subagent can read many files or long test output and return only a summary to the parent agent, so the parent's window stays smaller.
Can too much context make an agent worse?
Too much context can make an agent worse, because recall drops as the window fills. Long instruction files are also followed less reliably. Untrusted text in the window can carry prompt injection, so more untrusted input brings more risk.