Measure actual usage
System prompts, retrieved documents, tool definitions and conversation history all consume input tokens. Sample real requests rather than estimating from visible message length.
From a single request to a month of traffic. Estimate your LLM API bill using separate input, output, and cached token pricing.
No cache-read rate is available for this model. Cache savings are not assumed.
Apply input, output and cache-read pricing separately, then multiply by your request volume. All rates below are per 1 million tokens.
Each request uses 1,000 input and 500 output tokens. At $2 input, $8 output and $0.50 cached input per million: $12 uncached input + $2 cached input + $40 output.
$54.00 per monthIllustrative pricing, not a quote for a particular model.
System prompts, retrieved documents, tool definitions and conversation history all consume input tokens. Sample real requests rather than estimating from visible message length.
For models that bill reasoning tokens, include them in output usage even when they are not visible in the final answer. Check the provider’s usage breakdown.
Prompt caches usually have eligibility rules, minimum lengths and expiration times. Use an observed cache hit rate; 100% caching is rarely a safe starting assumption.