Context Window Budget Calculator
Give every token a job before the request is built. Allocate your window across system prompt, tools, retrieval, and history, then see what each call costs. The practice behind Context Engineering Is Not Prompt Engineering with a Better Name. Runs entirely in your browser.
What a call costs
Reading the bar
- Leave headroom. Past roughly 60 to 70 percent of the window, most models get measurably worse at using what you gave them. A full window is a bug, not a flex.
- History is the silent killer. System prompt and tools are fixed costs you set once; conversation history grows every turn until it eats the budget. Decide your truncation or summarization policy now, not in the incident review.
- The surcharge cliff is real. OpenAI bills 2x input past 272K tokens in a single request, Gemini Pro past 200K. The calculator flags when your budget crosses it.
Why deciding what the model never sees is most of the job: the essay.