Pricing
Per token. Nothing else.
No subscription, no seats, no minimum and no commitment. Add credit, spend it per token, add more when you want to.
| Model | Input / 1M | Output / 1M | Cache read / 1M | Context | |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash deepseek/deepseek-v4.1-flash50% off | $0.300$0.150 $0.090 off-peak | $1.200$0.600 $0.360 off-peak | $0.006 $0.003 off-peak | 1049K | Use this model |
| DeepSeek V4 Pro deepseek/deepseek-v4-pro50% off | $1.320$0.660 | $3.960$1.980 | $0.044 | 1049K | Use this model |
| DeepSeek V4 Flash deepseek/deepseek-v4-flash50% off | $0.440$0.220 | $1.320$0.660 | $0.014 | 1049K | Use this model |
| GLM 5.3 Flash z-ai/glm-5.3-flash | $0.150 | $0.500 | $0.030 | 1049K | Use this model |
| GLM 5.2 z-ai/glm-5.250% off | $1.400$0.700 | $4.400$2.200 | $0.260 | 1049K | Use this model |
- Input / 1M
- $0.300$0.150$0.090 off-peak
- Output / 1M
- $1.200$0.600$0.360 off-peak
- Cache read / 1M
- $0.006$0.003 off-peak
- Context
- 1049K
- Input / 1M
- $1.320$0.660
- Output / 1M
- $3.960$1.980
- Cache read / 1M
- $0.044
- Context
- 1049K
- Input / 1M
- $0.440$0.220
- Output / 1M
- $1.320$0.660
- Cache read / 1M
- $0.014
- Context
- 1049K
- Input / 1M
- $0.150
- Output / 1M
- $0.500
- Cache read / 1M
- $0.030
- Context
- 1049K
- Input / 1M
- $1.400$0.700
- Output / 1M
- $4.400$2.200
- Cache read / 1M
- $0.260
- Context
- 1049K
Struck-through figures are the model vendor’s own published rate; the price beside them is what we charge. Billed per token, no minimum. Cache reads are the exception and carry no discount — they already cost between a third and a twenty-fifth of the input rate, and we pass the supplier’s cache rate through rather than flattening it into one headline number.
Off-peak pricing. DeepSeek V4.1 Flash is billed at two rates. Peak hours are Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8); everything else is off-peak, including weekends. A request is billed at the rate in force when it is made, and every request in your logs records which one it was.
The listed discount applies to the first token and the hundred-millionth alike. Credit never expires and unspent credit is refundable.
Comparison
What the same traffic costs at DeepSeek and Zhipu list price
Blended at a 3:1 prompt-to-completion ratio, which is roughly what production traffic looks like.
| Model | Input | Output | Cache read | API format |
|---|---|---|---|---|
| GPT-6 LunaOpenAI | $0.100 | $0.500 | $0.010 | OpenAI |
| Claude Haiku 4.5Anthropic | $1.000 | $5.000 | $0.100 | Anthropic |
| Gemini 3.8 FlashGoogle | $0.750 | $3.750 | $0.075 | |
| DeepSeek V4.1 FlashTokenify, peak | $0.150 | $0.600 | $0.006 | OpenAI and Anthropic |
| DeepSeek V4.1 FlashTokenify, off-peak | $0.090 | $0.360 | $0.003 | OpenAI and Anthropic |
Per 1M tokens. Our rows come from the live catalogue at build time. Other vendors’ prices were read from their own pages and last checked 2026-10-01: openai.com (2026-09-30), claude.com (2026-10-01), ai.google.dev (2026-10-01).
Gemini 3.8 Flash is on a promotional rate until 2026-12-31; it becomes $1.50 / $7.50 / $0.15 on 1 January 2027.
Time of day
Two rates on the models whose vendor charges two rates
DeepSeek V4.1 Flash is sold at a peak rate and a cheaper off-peak one, because that is how the model vendor sells it. Peak is Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8); every other hour, including the whole weekend, is off-peak.
Which rate you pay is decided when the request arrives, not when it finishes. A generation that starts at 17:59 and runs for four minutes is billed entirely at the peak rate, and one that starts a minute later is billed entirely at the off-peak rate. Nothing is split across the boundary, so a long request can never cost more than the rate you sent it at.
Every request records which window it was billed in, and it is on your own activity page — so a batch job moved into the cheap hours can be checked rather than taken on trust. Both rates and the window itself are on /v1/models as pricing, pricing_offpeak and pricing_window, so a client can price a request before making it.
Billing
How a request turns into a charge
You are billed per token against the usage the model vendor reports — the same numbers on your invoice line and in ours — with no rounding up to a minimum and no per-request fee.
Cached prompt tokens are billed separately, at the cache read rate rather than the input rate. Those already cost between a third and a twenty-fifth of a fresh input token depending on the model, and we pass that through rather than flattening it into one blended rate. On agent workloads with a long stable system prompt it is usually the largest saving of all.
Credit is reserved when a request starts and the unused part is refunded the moment it finishes, so a runaway generation stops at your balance instead of past it. Reasoning models are the exception worth knowing about: their reasoning is not bounded by max_tokens, so the reservation carries a measured allowance for it and an unusually long run can settle slightly above what was held.
A 4xx from upstream is never charged, and neither is a request that fails before any tokens are generated.
Last updated 2026-10-01. Prices come from the live catalogue at the time this page was built.