Here is the uncomfortable truth about most "prompt engineering" effort: the wording of your instructions is the smallest lever in the system. I have watched teams spend a sprint rewording a system prompt while the model was drowning in forty thousand tokens of stale tool output. The prompt was fine. The context was garbage.

Context engineering is the name that finally stuck for the real work: deciding what the model sees at inference time, in what order, at what level of detail, under what token budget. The term is new. The discipline is not. Anyone who has shipped a serious agent or retrieval system has been doing it for two years, usually badly, usually by accident.

This essay is my attempt to write down the rules I keep relearning.

What Context Engineering Actually Is

Context engineering is the practice of assembling everything a model sees for a given call: instructions, retrieved evidence, tool results, conversation history, and working state. It covers selection (what gets in), structure (how it is organized and labeled), budget (how many tokens each part deserves), order (what comes first and last), and freshness (when cached context must be rebuilt). Prompt engineering tunes the instructions. Context engineering governs everything around them, which in a production system is 95% of the window.

The distinction matters because the failure modes are different. A bad prompt gives you wrong behavior consistently, which is easy to catch. Bad context gives you wrong behavior intermittently, keyed to which documents got retrieved, how long the session has run, and what a tool dumped into the transcript three steps ago. Intermittent failures are the expensive kind.

Why the Term Showed Up Now

Two shifts made this a named discipline. First, agents. A single chat completion has one context to curate. An agent running twelve steps has a context that mutates on every step: tool outputs append, history grows, and the question the user actually asked drifts further from the model's attention. In multi-agent systems the problem compounds, because every agent boundary is a decision about what context crosses it.

Second, long context windows made curation matter more, not less. The marketing said you could stop choosing and just load everything. Practice said otherwise. Models attend unevenly across long windows, the middle gets lost, cost scales with every token you did not need, and irrelevant context actively degrades answers. A million-token window is a bigger room to fill with junk. I wrote about the storage side of this tradeoff in RAG vs. CAG vs. KAG; context engineering is the discipline that sits on top of whichever storage answer you picked.

The Failure Modes That Actually Hurt

Context flooding. A tool returns 30KB of JSON and someone passes it through verbatim. Do this three times and the model is reasoning about your question through a keyhole. The fix is boring: every tool gets an output contract, and verbose results get truncated or summarized before they enter the window.

Context rot. Long-running sessions accumulate stale facts: an earlier answer that was corrected, a plan that was abandoned, a file that has since changed. The model cannot tell stale from current unless you mark it or remove it. Sessions need compaction the way data lakes need governance, and for the same reason: unmanaged accumulation turns an asset into sludge.

Distraction. Relevance is not binary. Ten marginally related documents reliably beat one perfect document into the wrong answer, because the model tries to honor everything you showed it. Cutting context is usually a bigger quality win than adding it.

Starving the question. The user's actual request ends up as fifty tokens at the bottom of a hundred-thousand-token sandwich. Position is a resource. The task belongs where attention is strongest: stated up front, restated at the end if the window is long.

The Patterns That Work

Documents Tool results Memory / state History Select relevance, recency Compact summarize, dedupe Context window system + task (pinned) curated evidence working state the actual question budgeted and ordered
Context assembly is a pipeline with a budget, not a string concatenation.

Budget like an engineer. Give every section of the window an explicit token allocation and enforce it in code. When the evidence section overflows, the selector tightens or the compactor kicks in. Nothing enters unbudgeted. This one habit eliminates most flooding incidents, and it is ten lines of code:

# Context assembly with hard budgets (simplified from production)
BUDGET = {
    "system":    1_500,   # pinned, never trimmed
    "task":        800,   # pinned, never trimmed
    "evidence": 12_000,   # retrieval results, ranked
    "state":     4_000,   # working memory, compacted
    "history":   6_000,   # recent turns, oldest summarized first
}

def assemble(parts):
    window = []
    for section, limit in BUDGET.items():
        content = parts[section]
        if tokens(content) > limit:
            content = compact(content, limit)   # rank, trim, summarize
        window.append(labeled(section, content))
    return order_for_attention(window)  # task first and last

# The log line that saves you at 2am:
log.info("context", {s: tokens(parts[s]) for s in BUDGET})

Structure and label everything. Models follow sections better than soup. Every block in the window carries a label and, for evidence, provenance: where it came from and when. Labeled context is also the only context you can debug later.

Prefer just-in-time over preload. Agents that fetch what a step needs, when it needs it, beat agents that carry everything from step one. The window stays small, the evidence stays fresh, and the retrieval itself can reason about what is missing.

Separate working memory from the transcript. Long tasks need a scratchpad the agent maintains deliberately: current plan, decisions made, facts confirmed. Compact the transcript aggressively and keep the scratchpad authoritative. When the two disagree, the scratchpad wins, because the transcript is where rot lives.

Make context observable. You cannot fix what you cannot see. Log the assembled window, or at minimum its shape: sections, token counts, provenance. The single most useful debugging question in applied AI is no longer "what did the model say." It is "what did the model see."

I keep watching the same scene in production reviews. An agent fails a task, the team reads the final answer and starts theorizing about model quality, and then someone finally dumps the assembled context. The answer is almost always sitting right there: the critical document got ranked out of the evidence section, or a tool response ate the budget, or the instruction that mattered scrolled into the summarized zone. The model did fine with what it was shown. What it was shown was the bug.

Where This Meets Evals

Context engineering without measurement is vibes with extra steps. The eval suite that catches context bugs scores the window, not just the answer: was the needed evidence present, what fraction of tokens were relevant, did the task survive compaction. If you have the eval flywheel running, add context-level assertions to it. A regression that drops recall in your evidence selector will look exactly like "the model got dumber" until something is measuring the window.

The Bottom Line

Prompt engineering had a good run as a job title, but the durable skill was never the wording. It is the information architecture of the model's working set: what gets in, what gets cut, what order, what budget, and how you see it when it breaks. That is engineering in the plain sense of the word, with budgets and contracts and logs.

The teams whose agents feel smart are not running better models. They are showing their models better context.

Context Engineering: Direct Answers

What is context engineering?
The practice of assembling what a model sees at inference time: selecting evidence, structuring and labeling it, enforcing token budgets per section, ordering for attention, and keeping it fresh across steps. It spans retrieval, tool outputs, memory, and conversation history.

How is context engineering different from prompt engineering?
Prompt engineering tunes the instructions. Context engineering governs everything else in the window, which in production systems is the overwhelming majority of it. Prompts fail consistently; context fails intermittently, which makes it the harder and more valuable problem.

Does long context make context engineering unnecessary?
The opposite. Large windows raise the cost of careless assembly: attention degrades across long contexts, irrelevant material actively hurts answers, and every unneeded token is billed. A bigger window is a bigger thing to curate.

What is agentic context engineering?
The same discipline applied across multi-step agents, where context mutates every step: tool-output contracts, scratchpad memory separate from the transcript, compaction between steps, and deciding what context crosses each agent boundary.

How do you debug context problems?
Log what the model saw. Reconstruct the assembled window for the failing call, check whether the needed evidence was present and where it sat, and add a context-level assertion to your eval suite so the regression cannot return silently.