All model pricing

Token pricing calculator

From a single request to a month of traffic. Estimate your LLM API bill using separate input, output, and cached token pricing.

01

Choose model pricing or enter custom rates

Input $/1M $4.40Output $/1M $22.00
02

Tell us about your workload

Input tokens include cache hits. Cached tokens are charged once, at the cache-read rate.

How token pricing becomes your API cost

Apply input, output and cache-read pricing separately, then multiply by your request volume. All rates below are per 1 million tokens.

Monthly cost = requests × [(uncached input × input rate + cached input × cache rate + output × output rate) ÷ 1,000,000]

10,000 requests. 40% cache hits.

Each request uses 1,000 input and 500 output tokens. At $2 input, $8 output and $0.50 cached input per million: $12 uncached input + $2 cached input + $40 output.

$54.00 per month

Illustrative pricing, not a quote for a particular model.

Build a budget you can trust

Measure actual usage

System prompts, retrieved documents, tool definitions and conversation history all consume input tokens. Sample real requests rather than estimating from visible message length.

Include reasoning output

For models that bill reasoning tokens, include them in output usage even when they are not visible in the final answer. Check the provider’s usage breakdown.

Model realistic cache hits

Prompt caches usually have eligibility rules, minimum lengths and expiration times. Use an observed cache hit rate; 100% caching is rarely a safe starting assumption.