LLM API Cost Calculator
Price your actual workload, not a marketing table. Tokens per request, daily volume, cache hit rate, and batch discounts across Claude, GPT, and Gemini. Every price is editable. Runs entirely in your browser.
0% of input cached
Cached input is billed at roughly 10% of the input rate across providers; the slider applies that to the cached share. Long-context surcharges apply automatically when input tokens per request cross a provider's threshold. Prices are standard tier per 1M tokens, verified , and editable below if they have drifted.
| Model | $/1M in · out | Per request | Per day | Per month |
|---|
How to read this honestly
- Cheapest is not best. This table answers one question: what the same traffic costs per provider. Whether the cheap model clears your quality bar is a question for your eval set, and that is the order to do it in: quality gate first, then this table.
- Cache hit rate is an engineering outcome. Stable system prompts and tool definitions at the front of the request cache beautifully; anything after a dynamic value does not. A workload redesign that moves you from 0 to 80 percent cached often beats switching providers.
- Batch when latency does not matter. Overnight summarization, backfills, and eval runs have no business paying real-time rates.
- Model prices change monthly. If the verified date above is stale, edit the price cells; every number recomputes.
For the sizing side of the same conversation, pair this with the Context Window Budget Calculator. For why most AI project budgets die anyway: The $400B Lie.