

DeepSeek
DeepSeek pricing charges per million tokens and splits input by whether the token hit the disk cache.
Updated on:
DeepSeek pricing: the cache decides what an input token costs
DeepSeek pricing charges per million tokens and splits input by whether the token hit the disk cache. A cache hit on deepseek-flash costs $0.003 per million off-peak, a cache miss costs $0.15, and that's a 50x gap on the same token. Every rate then doubles during peak hours. Nothing else in this index prices the lookup outcome rather than the act of caching, so what you pay depends on something you don't directly control.
Key takeaways
DeepSeek charges cache-hit and cache-miss input at separate rates, 50x apart on deepseek-flash and 30x apart on deepseek-v4-pro.
Peak hours run 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, and double every rate. That's 35 hours of a 168-hour week, so off-peak applies roughly 79% of the time.
Caching is on by default with no parameter to set, and DeepSeek charges nothing to write a cache entry.
There's no plan, no allowance and no overage. Calls return HTTP 402 the moment the prepaid balance hits zero.
DeepSeek pricing in 2026
Model | Cache hit input | Cache miss input | Output | Concurrency
|
|---|---|---|---|---|
deepseek-flash (V4.1 Flash) | $0.003 / $0.006 | $0.15 / $0.30 | $0.60 / $1.20 | 2,500 |
deepseek-v4-pro (V4-Pro-0813) | $0.022 / $0.044 | $0.66 / $1.32 | $1.98 / $3.96 | 500 |
Source: api-docs.deepseek.com/quick_start/pricing, read 22 September 2026.
What DeepSeek actually meters
DeepSeek meters tokens, and it's the only vendor in this index that prices the outcome of a cache lookup rather than the work of building the cache. OpenAI and Anthropic both charge a premium to write a cache entry and a discount to read one. DeepSeek charges nothing to write and applies one of two input rates depending on whether your prefix matched.
Every call returns prompt_cache_hit_tokens and prompt_cache_miss_tokens in its usage block, so the split is auditable per request rather than inferred from the invoice.
The gap is severe enough to dominate a bill. On deepseek-flash a cache hit costs 2% of a cache miss. On deepseek-v4-pro it costs roughly 3.3%. Output is the milder multiple: 4x cache-miss input on Flash and 3x on Pro, a narrower spread than the 5x most Western providers run.
Caching itself is automatic, enabled for all users by default with no parameter to toggle. Hits require a full match against a persisted prefix unit, and DeepSeek persists those at request boundaries, at common prefixes detected across requests, and at fixed token intervals inside long inputs. The docs are explicit that this is best effort with no guaranteed hit rate, and that an unused cache clears within hours to days.
How DeepSeek's balance works
DeepSeek runs prepaid, with two balances per account. A granted balance covers promotional credit and a topped-up balance covers what you've paid in. Charges draw from the granted balance first when both exist, which is the deduction order you'd want.
Granted credit expires. The balance endpoint returns granted_balance as "the total not expired granted balance", which confirms expiry exists, but the docs never publish the window or what triggers a grant. That's the one real gap in an otherwise complete rate card.
There are no credit packs, rollover rules or pooling to document, because DeepSeek doesn't sell credits as a product. The balance is a cash float against a meter. A user_id parameter isolates content safety, cache and scheduling per end user, but it doesn't split the balance.
What happens when you hit the limit
Nothing throttles and nothing auto-charges. At zero balance the API returns HTTP 402 Insufficient Balance and you top up to resume, which is a hard stop rather than an invoice surprise.
The separate ceiling is concurrency: 2,500 in-flight requests on Flash and 500 on Pro, counted per account across all API keys. Exceeding it returns HTTP 429. DeepSeek grants expansion on request and states plainly that it costs nothing extra, which is unusual. Most providers attach higher concurrency to a spend tier.
How DeepSeek pricing has changed across all these years
Date | Milestone | Source
|
|---|---|---|
10 Sep 2026 | V4.1 Flash ships and cuts Flash rates: cache miss $0.22 to $0.15 off-peak, cache hit $0.007 to $0.003, output $0.66 to $0.60. Pro unchanged | Vendor |
16 Aug 2026 | Flash cache-miss input moves from a flat $0.14 to $0.22 off-peak and $0.44 peak. Pro moves from $0.435 to $0.66 and $1.32 | Vendor |
13 Aug 2026 | Peak and off-peak billing announced, off-peak set at half of peak, effective 16:00 UTC on 16 Aug 2026 | Vendor |
24 Apr 2026 | V4-Pro and V4-Flash replace deepseek-chat and deepseek-reasoner on the API | Vendor |
21 Aug 2025 | New rate card announced with off-peak discounts ending 5 Sep 2025 at 16:00 UTC | Vendor |
Customer
Sentiment Highlights
This notice is making me nervous
DeepSeek API user reacting to the announced price rise, Hacker News, August 2026
Explore other providers

Deepgram
AI Voice
Deepgram pricing meters audio by the second and refuses to round.

Claude API
API
Claude API pricing charges per million tokens, and it splits them five ways.

Azure OpenAI
Enterprise LLM
Azure OpenAI pricing is the only entry in this index where you can pay for a model without sending it anything.
How much does DeepSeek cost per million tokens?
What are DeepSeek's peak hours?
Do I need to enable DeepSeek prompt caching?
Is there a free tier on the DeepSeek API?
























