DeepSeek API pricing, compared
What DeepSeek charges, what the resellers charge, and what a month of real traffic costs at each. Third-party prices were read from their own pages and are dated; ours come from the live catalogue.
In short
Three things decide the bill
- DeepSeek prices by the clock. Peak is exactly twice off-peak, and peak is Mon–Fri 01:00–04:00 and 06:00–10:00 UTC — a weekday window, not nights and weekends.
- We charge the off-peak rate as our peak rate. DeepSeek V4.1 Flash is $0.150 per 1M input here at any hour, and $0.090 in our own off-peak window.
- Cache reads are where an agent’s money goes. Nine tokens in ten on an agent workload are cache reads, so the cache rate moves the monthly bill more than the headline input rate does.
The vendor
DeepSeek's own prices
Both windows, from their pricing page.
| Model | Window | Input | Output | Cache read |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | peak | $0.300 | $1.20 | $0.006 |
| DeepSeek V4.1 Flash | offpeak | $0.150 | $0.600 | $0.003 |
| DeepSeek V4 Pro | peak | $1.32 | $3.96 | $0.044 |
| DeepSeek V4 Pro | offpeak | $0.660 | $1.98 | $0.022 |
Peak is Mon–Fri 01:00–04:00 and 06:00–10:00 UTC, excluding Chinese public holidays. Read from api-docs.deepseek.com on 2026-10-01.
The sellers
The same model, through different sellers
One table per model. Ours is the highlighted row.
DeepSeek V4.1 Flash
| Seller | Input | Output | Cache read | Notes and source |
|---|---|---|---|---|
| DeepSeek (official) (peak) | $0.300 | $1.20 | $0.006 | api-docs.deepseek.com · 2026-10-01 |
| DeepSeek (official) (offpeak) | $0.150 | $0.600 | $0.003 | api-docs.deepseek.com · 2026-10-01 |
| Together | $0.300 | $1.20 | $0.006 | www.together.ai · 2026-10-01 |
| SiliconFlow | $0.300 | $1.20 | $0.006 | fp8openrouter.ai · 2026-10-01 |
| Novita | $0.300 | $1.20 | $0.006 | fp8openrouter.ai · 2026-10-01 |
| Tokenify (peak) | $0.150 | $0.600 | $0.006 | OpenAI and Anthropic formats; card or USDT/USDC; credit does not expire |
| Tokenify (off-peak) | $0.090 | $0.360 | $0.003 | our own off-peak window |
DeepSeek V4 Pro
| Seller | Input | Output | Cache read | Notes and source |
|---|---|---|---|---|
| DeepSeek (official) (peak) | $1.32 | $3.96 | $0.044 | api-docs.deepseek.com · 2026-10-01 |
| DeepSeek (official) (offpeak) | $0.660 | $1.98 | $0.022 | api-docs.deepseek.com · 2026-10-01 |
| Together | $1.32 | $3.96 | $0.130 | www.together.ai · 2026-10-01 |
| SiliconFlow | $1.32 | $3.96 | $0.044 | fp8openrouter.ai · 2026-10-01 |
| Novita | $0.990 | $2.97 | $0.033 | fp8openrouter.ai · 2026-10-01 |
| OpenRouter | $0.660 | $1.98 | $0.022 | default route; a marketplace, so the provider behind it can changeopenrouter.ai · 2026-10-01 |
| DeepInfra | $1.30 | $2.60 | $0.100 | fp8openrouter.ai · 2026-10-01 |
| Tokenify (peak) | $0.660 | $1.98 | $0.044 | OpenAI and Anthropic formats; card or USDT/USDC; credit does not expire |
This lists the sellers we have chosen to be compared against, not every seller. Some charge less than we do, and on these two models every one of them that publishes a precision is serving a quantised copy — fp8 or fp4 rather than the vendor’s own weights. That is the trade those prices represent, and it is worth knowing before reading a cheaper number as the same product. We are not claiming to be cheapest anywhere; these are the sellers shown, with the date each figure was read.
A month of traffic
What that works out to
Three workloads, costed at each seller's published rates. Token counts are assumptions and are stated; the prices are not.
Coding agent
An agent resends its system prompt, its tool schemas and every file it has read on each turn, so most of the prompt is a cache read. Work happens in office hours, which is when the vendor charges most. 100M input, 90% of it cached, 2M output per month.
| Seller | DeepSeek V4.1 Flash | DeepSeek V4 Pro |
|---|---|---|
| DeepSeek (official) | $5.94 | $25.08 |
| Together | $5.94 | $32.82 |
| SiliconFlow | $5.94 | $25.08 |
| Novita | $5.94 | $18.81 |
| OpenRouter | — | $12.54 |
| DeepInfra | — | $27.20 |
| Tokenify | $3.24 | $14.52 |
Overnight batch
Classification, extraction or summarisation run on a schedule. Nothing is waiting on it, so all of it can run in the cheap window. 500M input, 50M output per month.
| Seller | DeepSeek V4.1 Flash | DeepSeek V4 Pro |
|---|---|---|
| DeepSeek (official) | $105 | $429 |
| Together | $210 | $858 |
| SiliconFlow | $210 | $858 |
| Novita | $210 | $644 |
| OpenRouter | — | $429 |
| DeepInfra | — | $780 |
| Tokenify | $63.00 | $429 |
Product traffic
Requests from users, arriving around the clock. 21% land in the peak window because that is the share of the week it covers. 100M input, 10M output per month.
| Seller | DeepSeek V4.1 Flash | DeepSeek V4 Pro |
|---|---|---|
| DeepSeek (official) | $25.37 | $104 |
| Together | $42.00 | $172 |
| SiliconFlow | $42.00 | $172 |
| Novita | $42.00 | $129 |
| OpenRouter | — | $85.80 |
| DeepInfra | — | $156 |
| Tokenify | $14.35 | $85.80 |
Our side
How these prices are set
Every model has a published list price set by the company that made it. Ours is that list price less a fixed discount, which is why the saving is the same on input and on output and does not depend on how much you buy — there is no volume tier to reach and no negotiated rate.
Cache reads are the exception and are not discounted. They already cost a fraction of a fresh token, and the traffic that leans on them is the traffic worth having.
Prices on this site are read from the live catalogue when the page is built, so what you see here is what the API will charge. A price change is recorded and shown in the dashboard against the requests it applied to.
Questions
Frequently asked
When are DeepSeek's peak hours?
Mon–Fri 01:00–04:00 and 06:00–10:00 UTC, excluding Chinese public holidays. Every other hour is off-peak and costs half as much. It is a short weekday window rather than the nights-and-weekends arrangement the phrase usually suggests, and it falls across an Asian working day: a service answering users in Europe or the Americas overlaps it barely or not at all, so most of that traffic is already off-peak.
Why is cache pricing not discounted?
A cached token is billed at DeepSeek’s published cache rate, which is already a small fraction of a fresh one. It is the number to compare if your workload has a long stable prompt: an agent sends nine cache reads for every fresh token, so the cache rate moves a monthly bill more than the headline input rate does.
How much cheaper is Tokenify than the official API?
DeepSeek V4.1 Flash is $0.150 per 1M input here against $0.300 at DeepSeek's peak rate, 50% less. DeepSeek V4 Pro is $0.660 per 1M input here against $1.32 at DeepSeek's peak rate, 50% less.
Does the comparison list every seller?
No, and it does not claim to. It lists the sellers we have chosen to be compared against, each with the page its price came from and the date it was read. There are hosts quantising these models to fp8 or fp4 and charging less than we do; a table built to make us look best would leave them out quietly rather than say so here.
Which API formats work?
Both. The same key and the same models answer on the OpenAI chat-completions format and on the Anthropic Messages format, so a client written for either runs unchanged by pointing it at a different base URL.
DeepSeek V4.1 Flash · DeepSeek V4 Pro · our full rate card · getting a key
Last updated 2026-10-01. Prices come from the live catalogue at the time this page was built.