LLM Model Reference
Every production model that matters, one table: API price, context window, max output, and the long-context surcharge cliffs the pricing pages bury. I verify this by hand and date every check, because a stale pricing table is worse than none.
| Model | $/1M input | $/1M output | Context | Max out | Surcharge |
|---|
Standard tier. Batch APIs run about 50% off at every provider; cached input bills at roughly 10% of the input rate. To price your own workload with these numbers, use the LLM API Cost Calculator; to plan a window, the Context Window Budget Calculator.
Changelog
Spotted a drift before I did? Tell me on LinkedIn and I will verify and log it.
How to read the surcharge column
- OpenAI: a single request with more than 272K input tokens bills the entire request at 2x input and 1.5x output. Not the overage. The whole request.
- Gemini Pro: same shape past 200K input. The Flash tiers stay flat.
- Anthropic: no surcharge tier on the current lineup at standard context.
- If your retrieval budget hovers near a threshold, the surcharge is a bigger cost lever than the model choice. The budget calculator flags it automatically.