Token pricing units
Pricing is in USD per 1 million tokens for standard, global, default-group API usage. Input, output and cache reads stay separate. Missing pricing remains unknown; a reported zero stays zero.
Compare token pricing across providers, explore new models, and estimate your API bill.
Compare provider pricing for models with the highest recent token usage.
DeepSeek · Listed Sep 10, 2026 · 1M context
26 providers
Z.ai · Listed Aug 26, 2026 · 1.3M context
33 providers
DeepSeek · Listed Jul 31, 2026 · 1.3M context
28 providers
DeepSeek · Listed Apr 24, 2026 · 1M context
17 providers
22 models · 79 offersExpand a row to compare provider pricing
The scheduled data check is overdue. These are the last successfully published rates; confirm the source before making a purchase.
Scroll to compare all pricing →
Pricing is in USD per 1 million tokens for standard, global, default-group API usage. Input, output and cache reads stay separate. Missing pricing remains unknown; a reported zero stays zero.
Each offer links to the page where its pricing was published. The first row is that model’s base rate. Other rows are additional provider pricing published by LMSpeed. A published listing is not an independent endorsement. Only active, standard base token rates are included.
This table excludes batch, region-specific, VIP-group and token-tier offers. Taxes, cache writes, storage, image/audio billing and per-request tools may add charges. Capabilities are reported metadata, not a performance benchmark.
Understand token pricing before estimating your API bill.
A token is a unit of text processed by a model: it may be a word, part of a word, punctuation or whitespace. Tokenization varies by model and language. Use your provider’s tokenizer or the usage reported by real API requests for budgeting.
To convert token pricing, multiply the per-token rate by 1,000,000. A rate of $0.000002 per input token is $2 per million input tokens. Divide a per-million rate by 1,000,000 to go the other way. Input and output often have different rates.
The cached column shows pricing for reading eligible input tokens from a provider’s prompt cache. A cache hit replaces the normal input charge for those tokens. Cache creation, retention and storage can carry separate charges and are not included in that column.
The catalog is checked daily when synchronization is configured. The visible timestamp is the source data’s update time; a successful check with no changes does not reset it. If checks fail, the previous snapshot remains available and an overdue notice appears.
No. Long answers can make output pricing dominate the bill, while repeated context can favor a model with prompt caching. Compare your input/output mix and required capabilities in the calculator. Quality, latency, limits and tier conditions still matter.