Tokenify

Pricing

Per token. Nothing else.

No subscription, no seats, no minimum and no commitment. Add credit, spend it per token, add more when you want to.

DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash
50% off
Input / 1M
$0.300$0.150
$0.090 off-peak
Output / 1M
$1.200$0.600
$0.360 off-peak
Cache read / 1M
$0.006
$0.003 off-peak
Context
1049K
Use this model
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
50% off
Input / 1M
$1.320$0.660
Output / 1M
$3.960$1.980
Cache read / 1M
$0.044
Context
1049K
Use this model
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
50% off
Input / 1M
$0.440$0.220
Output / 1M
$1.320$0.660
Cache read / 1M
$0.014
Context
1049K
Use this model
GLM 5.3 Flash
z-ai/glm-5.3-flash
Input / 1M
$0.150
Output / 1M
$0.500
Cache read / 1M
$0.030
Context
1049K
Use this model
GLM 5.2
z-ai/glm-5.2
50% off
Input / 1M
$1.400$0.700
Output / 1M
$4.400$2.200
Cache read / 1M
$0.260
Context
1049K
Use this model

Struck-through figures are the model vendor’s own published rate; the price beside them is what we charge. Billed per token, no minimum. Cache reads are the exception and carry no discount — they already cost between a third and a twenty-fifth of the input rate, and we pass the supplier’s cache rate through rather than flattening it into one headline number.

Off-peak pricing. DeepSeek V4.1 Flash is billed at two rates. Peak hours are Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8); everything else is off-peak, including weekends. A request is billed at the rate in force when it is made, and every request in your logs records which one it was.

No tiers, no minimum, no expiry

The listed discount applies to the first token and the hundred-millionth alike. Credit never expires and unspent credit is refundable.

Comparison

What the same traffic costs at DeepSeek and Zhipu list price

Blended at a 3:1 prompt-to-completion ratio, which is roughly what production traffic looks like.

ModelInputOutputCache readAPI format
GPT-6 LunaOpenAI$0.100$0.500$0.010OpenAI
Claude Haiku 4.5Anthropic$1.000$5.000$0.100Anthropic
Gemini 3.8 FlashGoogle$0.750$3.750$0.075Google
DeepSeek V4.1 FlashTokenify, peak$0.150$0.600$0.006OpenAI and Anthropic
DeepSeek V4.1 FlashTokenify, off-peak$0.090$0.360$0.003OpenAI and Anthropic

Per 1M tokens. Our rows come from the live catalogue at build time. Other vendors’ prices were read from their own pages and last checked 2026-10-01: openai.com (2026-09-30), claude.com (2026-10-01), ai.google.dev (2026-10-01).

Gemini 3.8 Flash is on a promotional rate until 2026-12-31; it becomes $1.50 / $7.50 / $0.15 on 1 January 2027.

Time of day

Two rates on the models whose vendor charges two rates

DeepSeek V4.1 Flash is sold at a peak rate and a cheaper off-peak one, because that is how the model vendor sells it. Peak is Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8); every other hour, including the whole weekend, is off-peak.

Which rate you pay is decided when the request arrives, not when it finishes. A generation that starts at 17:59 and runs for four minutes is billed entirely at the peak rate, and one that starts a minute later is billed entirely at the off-peak rate. Nothing is split across the boundary, so a long request can never cost more than the rate you sent it at.

Every request records which window it was billed in, and it is on your own activity page — so a batch job moved into the cheap hours can be checked rather than taken on trust. Both rates and the window itself are on /v1/models as pricing, pricing_offpeak and pricing_window, so a client can price a request before making it.

Billing

How a request turns into a charge

You are billed per token against the usage the model vendor reports — the same numbers on your invoice line and in ours — with no rounding up to a minimum and no per-request fee.

Cached prompt tokens are billed separately, at the cache read rate rather than the input rate. Those already cost between a third and a twenty-fifth of a fresh input token depending on the model, and we pass that through rather than flattening it into one blended rate. On agent workloads with a long stable system prompt it is usually the largest saving of all.

Credit is reserved when a request starts and the unused part is refunded the moment it finishes, so a runaway generation stops at your balance instead of past it. Reasoning models are the exception worth knowing about: their reasoning is not bounded by max_tokens, so the reservation carries a measured allowance for it and an unusually long run can settle slightly above what was held.

A 4xx from upstream is never charged, and neither is a request that fails before any tokens are generated.

Last updated 2026-10-01. Prices come from the live catalogue at the time this page was built.