- Type
- Engineering practice
- Also called
- Context curation, attention management
- Emerged
- 2025
- Builds on
- Prompt engineering
- Key constraints
- Finite context window, attention budget, context rot
- Related
- RAG, agents, prompt engineering, MCP
- Type
- Engineering practice
- Also called
- Context curation, attention management
- Emerged
- 2025
- Builds on
- Prompt engineering
- Key constraints
- Finite context window, attention budget, context rot
- Related
- RAG, agents, prompt engineering, MCP
Background
For most of the generative AI era the dominant skill was prompt engineering — writing effective system prompts for single-turn classification or generation tasks. As applications moved from one-shot calls to agents running loops over many turns, the prompt became only one part of a much larger state: tool schemas, MCP server descriptions, retrieved files, scratchpads, prior messages and external data all land in the same window.
The term was popularised in mid-2025. LangChain described context engineering as "the most important skill an AI engineer can develop" in June 2025, arguing that most unreliable agents fail because the right context, instructions and tools were never communicated to the model.[2] Anthropic's engineering team published a detailed treatment in September 2025, framing context as a finite resource with diminishing marginal returns.[1]
Key Concepts
Context as a finite resource. Every token added to the window consumes part of the model's attention budget. In the transformer architecture each token attends to every other token, producing n² pairwise relationships for n tokens; as sequences grow, retrieval precision and long-range reasoning degrade.[1] Research on context rot has shown empirically that recall of information placed in a long context falls as token count rises, across all model families.[3]
Signal density. The working goal is the smallest possible set of high-signal tokens that maximises the chance of the desired outcome.[1] Practical techniques include:
- Compaction — summarising long histories into a shorter state so conversational flow continues without carrying every message.
- Note-taking — persisting durable facts to a file the agent can re-read, rather than holding everything in the window.
- Just-in-time retrieval — letting the agent search its environment when needed instead of pre-loading large corpora.
- Token-efficient tools — returning summaries rather than raw payloads, and truncating aggressively.
Progressive disclosure. Load only metadata at startup, instructions on activation, and detailed resources on demand — the same principle behind Agent Skills and tool descriptions.
Applications
Context engineering is the main lever behind reliable AI agents: coding assistants that retain project state across sessions, research agents that browse dozens of pages, and customer-support agents that must carry account context across a conversation. It also governs cost — tokens are billed on every call, so a bloated window inflates inference expense without improving output. Teams routinely find that trimming context improves both accuracy and latency, an outcome that inverts the "more context is better" assumption.
>Key Takeaways
- Context engineering curates the whole inference state, not just the system prompt.
- Context is finite: attention degrades as windows fill, a phenomenon called context rot.
- The guiding principle is the smallest high-signal token set that produces the desired behaviour.
- Compaction, note-taking and just-in-time retrieval are the core techniques for long-horizon agents.
- Better context usually lowers both hallucination rates and token spend.
See Also
🇲🇾 For Malaysian businesses deploying AI, context engineering is where reliability and cost are actually decided. A support bot for a local bank or telco must inject the right policy excerpt, customer record and language register into a limited window — too little and the model hallucinates, too much and answers degrade while the bill climbs in ringgit. Because models process context in tokens, Bahasa Malaysia and mixed-language (Manglish) inputs are not free: poor tokenisation of local languages can consume window budget faster than English, making careful truncation and retrieval strategy even more important. Firms following the PDPA must also treat retrieved context as personal data — dumping whole customer records into a prompt creates exposure that a filtered, purpose-limited extract does not. MDEC-aligned developers building on local and open-weight models, which often have smaller effective context windows, feel these constraints first.
References
- ↑Anthropic. Effective context engineering for AI agents (29 September 2025). https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- ↑LangChain. The rise of "context engineering" (23 June 2025). https://www.langchain.com/blog/the-rise-of-context-engineering
- ↑Chroma Research. Context Rot: How Long Context Impairs LLM Performance. https://research.trychroma.com/context-rot