AIWiki
Malaysia
Back to all articles
Applicationscontext windowagentsprompt engineering

Context Engineering

4 min readUpdated October 2026
Context Engineering
Type
Engineering practice
Also called
Context curation, attention management
Emerged
2025
Builds on
Prompt engineering
Key constraints
Finite context window, attention budget, context rot
Related
RAG, agents, prompt engineering, MCP
Context engineering is the discipline of choosing what information to place in a language model's context window — instructions, tool definitions, retrieved documents, conversation history — and keeping that information clean, relevant and compact across multi-step work. It is widely described as the successor to prompt engineering: rather than asking "what wording gets the best answer?", it asks "what state is the model in, and what is the smallest set of tokens that produces the behaviour we want?"[1]

Background

For most of the generative AI era the dominant skill was prompt engineering — writing effective system prompts for single-turn classification or generation tasks. As applications moved from one-shot calls to agents running loops over many turns, the prompt became only one part of a much larger state: tool schemas, MCP server descriptions, retrieved files, scratchpads, prior messages and external data all land in the same window.

The term was popularised in mid-2025. LangChain described context engineering as "the most important skill an AI engineer can develop" in June 2025, arguing that most unreliable agents fail because the right context, instructions and tools were never communicated to the model.[2] Anthropic's engineering team published a detailed treatment in September 2025, framing context as a finite resource with diminishing marginal returns.[1]

Key Concepts

Context as a finite resource. Every token added to the window consumes part of the model's attention budget. In the transformer architecture each token attends to every other token, producing n² pairwise relationships for n tokens; as sequences grow, retrieval precision and long-range reasoning degrade.[1] Research on context rot has shown empirically that recall of information placed in a long context falls as token count rises, across all model families.[3]

Signal density. The working goal is the smallest possible set of high-signal tokens that maximises the chance of the desired outcome.[1] Practical techniques include:

  • Compaction — summarising long histories into a shorter state so conversational flow continues without carrying every message.
  • Note-taking — persisting durable facts to a file the agent can re-read, rather than holding everything in the window.
  • Just-in-time retrieval — letting the agent search its environment when needed instead of pre-loading large corpora.
  • Token-efficient tools — returning summaries rather than raw payloads, and truncating aggressively.
Curated examples over instructions. Anthropic recommends replacing exhaustive rule lists with a small set of canonical examples that demonstrate the behaviour wanted, since examples constrain the model more reliably than prose.[1]

Progressive disclosure. Load only metadata at startup, instructions on activation, and detailed resources on demand — the same principle behind Agent Skills and tool descriptions.

Applications

Context engineering is the main lever behind reliable AI agents: coding assistants that retain project state across sessions, research agents that browse dozens of pages, and customer-support agents that must carry account context across a conversation. It also governs cost — tokens are billed on every call, so a bloated window inflates inference expense without improving output. Teams routinely find that trimming context improves both accuracy and latency, an outcome that inverts the "more context is better" assumption.

>Key Takeaways

  • Context engineering curates the whole inference state, not just the system prompt.
  • Context is finite: attention degrades as windows fill, a phenomenon called context rot.
  • The guiding principle is the smallest high-signal token set that produces the desired behaviour.
  • Compaction, note-taking and just-in-time retrieval are the core techniques for long-horizon agents.
  • Better context usually lowers both hallucination rates and token spend.

See Also

🇲🇾Malaysian Context

🇲🇾 For Malaysian businesses deploying AI, context engineering is where reliability and cost are actually decided. A support bot for a local bank or telco must inject the right policy excerpt, customer record and language register into a limited window — too little and the model hallucinates, too much and answers degrade while the bill climbs in ringgit. Because models process context in tokens, Bahasa Malaysia and mixed-language (Manglish) inputs are not free: poor tokenisation of local languages can consume window budget faster than English, making careful truncation and retrieval strategy even more important. Firms following the PDPA must also treat retrieved context as personal data — dumping whole customer records into a prompt creates exposure that a filtered, purpose-limited extract does not. MDEC-aligned developers building on local and open-weight models, which often have smaller effective context windows, feel these constraints first.

References

  1. ↑Anthropic. Effective context engineering for AI agents (29 September 2025). https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
  2. ↑LangChain. The rise of "context engineering" (23 June 2025). https://www.langchain.com/blog/the-rise-of-context-engineering
  3. ↑Chroma Research. Context Rot: How Long Context Impairs LLM Performance. https://research.trychroma.com/context-rot