Robots Atlas>ROBOTS ATLAS
Agents

Context Engineering

2025ActivePublished: 10 September 2026Updated: 10 September 2026Published
Key innovation
Shifts the focus from crafting a single prompt to continuously curating the entire contents of the context window (instructions, tools, history, retrieved data) as a finite resource on every inference cycle of an agent.
Category
Agents
Abstraction level
Paradigm
Operation level
Agent runtimeInferenceOrchestration
Use cases
Long-horizon AI agentsCoding agentsRAG systems and agentic searchMulti-step tool-using agent workflowsManaging memory and conversation history in chatbotsMulti-agent architectures (lead agent + sub-agents)

How it works

Context engineering combines several techniques applied on every inference cycle: (1) a system prompt at the right "altitude" — specific enough to steer behavior, flexible enough to leave the model heuristics; (2) a minimal, unambiguous tool set instead of bloated tool libraries; (3) diverse, canonical few-shot examples rather than exhaustive edge cases; (4) just-in-time retrieval — the agent keeps lightweight identifiers (file paths, URLs, queries) and loads data only when needed (progressive disclosure); (5) compaction — summarizing history before hitting the window limit while preserving architectural decisions and open threads; (6) structured note-taking / memory persisted outside the context window for long-horizon coherence; (7) sub-agent architectures with clean contexts that return condensed summaries (typically 1,000–2,000 tokens) to a coordinating lead agent.

Problem solved

It addresses the degradation of LLM-agent output quality caused by overloaded, poorly selected, or stale context. As the context window grows, "context rot" and the quadratic (O(n²)) attention cost increase and the model loses track of salient information; context engineering curates, prunes, and just-in-time delivers only the tokens that are needed.

Components

System promptSets the rules and operating frame for the agent on every inference cycle.

Persistent instructions defining agent behavior, written at the right "altitude" — specific enough to steer, general enough to leave the model heuristics.

Official

Tool definitionsLets the agent act on the external world without polluting the context.

A self-contained, unambiguous set of tools; bloated tool sets with overlapping functions should be avoided.

Official

Few-shot examplesCalibrates the model's output format and style.

Diverse, canonical examples illustrating expected behavior, rather than an exhaustive list of edge cases.

Official

Message historyProvides continuity and the agent's short-term memory.

The interaction so far, managed via curation strategies (pruning, compaction) to stay within the attention budget.

Retrieved dataDelivers only the facts needed at that moment (progressive disclosure).

Information loaded at runtime from lightweight identifiers (file paths, URLs, queries) instead of pre-loading all context upfront.

Official

External memory and notesMaintains agent knowledge beyond the limit of a single context window.

Durable notes and state persisted outside the context window, recalled on demand for long-horizon coherence.

Official

Implementation

Implementation pitfalls
Context rotHigh

Overfilling the context window degrades the model's ability to accurately recall information.

Fix:Keep context minimal and high-signal; apply compaction and just-in-time retrieval.
Tool bloatMedium

Too many overlapping tools increases token count and tool-selection errors.

Fix:Limit the set to self-contained, unambiguous tools.
Lost-in-the-middleHigh

Information placed in the middle of a long context is recalled less reliably than at the start and end.

Fix:Place key information at the context edges and trim ballast.
Stale context after compactionMedium

Aggressive summarization can drop details needed later (e.g. unresolved bugs, decisions).

Fix:Preserve architectural decisions and open threads; persist them to external memory.
Context poisoningHigh

Injecting erroneous or malicious content into context persists it across the agent's subsequent turns.

Fix:Validate and isolate retrieved data; limit trust in external content.

Evolution

Original paper · 2025 · Anthropic Engineering (blog), September 29, 2025 · Prithvi Rajasekaran
Effective context engineering for AI agents
Prithvi Rajasekaran, Ethan Dixon, Carly Ryan, Jeremy Hadfield
2020
Few-shot prompting in GPT-3

Placing examples in context as a way to steer the model — an early precursor of context curation.

2022
The prompt engineering era (ChatGPT)

Emphasis on wording a single prompt; the foundation that context engineering later generalized.

2023
RAG and context augmentation

Dynamically injecting retrieved knowledge into context popularized managing the window's contents.

2025
The term "context engineering" is formalized
Inflection point

The term was popularized in mid-2025 (Tobi Lütke, Andrej Karpathy), and LangChain described it in "The rise of context engineering" (June 23, 2025).

2025
Anthropic guide for AI agents
Inflection point

Anthropic formalized the practices (right-altitude system prompts, just-in-time retrieval, compaction, note-taking, sub-agents) in a guide published September 29, 2025.

Hyperparameters (configurable axes)

Context window budgetCritical

Target token count kept in context; a lower budget reduces context rot at the cost of less information.

8k–32k tokenówTypical range for task-oriented agents.
200k+ tokenówLarge windows require stronger curation.
Retrieval strategyHigh

Choice between pre-loading context and just-in-time retrieval.

just-in-timeLoad data on demand from lightweight identifiers.
upfront (RAG)Pre-fetch and inject chunks.
Compaction thresholdHigh

The window-fill level at which history summarization is triggered.

~80% oknaA common compaction trigger threshold.
Tool set sizeMedium

Number and granularity of tools exposed to the agent; too large a set causes bloat and selection errors.

5–15 narzędziUsually sufficient for a task agent.
Few-shot example countMedium

How many canonical examples to place in context to calibrate behavior without wasting token budget.

2–5 przykładówDiverse, canonical, not edge cases.
Memory strategyHigh

How state and notes are persisted outside the context window and the rules for recalling them.

structured note-takingNotes written to a file/external memory.
Sub-agent countMedium

How many specialized clean-context sub-agents a lead agent coordinates.

1–N sub-agentówEach returns a 1,000–2,000 token summary.

Computational complexity

Time complexity: O(n² · d). Space complexity: O(n · d · L).

Compute bottleneck

Long-context attention and KV-cache memory

As context length grows, attention cost (quadratic) and KV-cache footprint (linear) rise, limiting the model's effective "attention budget".

Depends on
Długość okna kontekstowegoLiczba warstw i głów uwagi

Execution paradigm

Primary mode
Conditional

Context contents are selected conditionally, depending on task state and query.

Activation pattern
Input dependent
Additional modes
Sparse
Routing mechanism

The agent decides at runtime which data to load and which parts of history to keep or summarize, instead of loading everything upfront.

Parallelism

Parallelism level
Partially parallel

Sub-agent architectures parallelize work, returning condensed summaries to a coordinating lead agent.

Scope
Inference
Constraints
!Context curation happens sequentially across the agent's inference cycles, but sub-agents can run in parallel with isolated contexts.

Hardware requirements

Primary

It is a software- and orchestration-level design practice, independent of any specific hardware.

Good fit

Reducing context length lowers KV-cache footprint and attention cost on GPUs during inference.