Tokenify

Models and pricing

Prices are per million tokens, billed per token with no rounding up to a minimum. Credit never expires.

The catalogue

Model idInputOutputCache readContext
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash50% off
$0.300$0.150$0.090 off-peak$1.200$0.600$0.360 off-peak$0.0060$0.0030 off-peak1024K
DeepSeek V4 Pro
deepseek/deepseek-v4-pro50% off
$1.320$0.660$3.960$1.980$0.0441024K
DeepSeek V4 Flash
deepseek/deepseek-v4-flash50% off
$0.440$0.220$1.320$0.660$0.0141024K
GLM 5.3 Flash
z-ai/glm-5.3-flash
$0.150$0.500$0.0301024K
GLM 5.2
z-ai/glm-5.250% off
$1.400$0.700$4.400$2.200$0.2601024K

Struck-through figures are the model vendor’s own published rate; the figure beside them is what we charge. This table is generated from the same catalogue the gateway routes on, so it cannot drift from what you are actually charged.

The figures in the table are the peak rates. DeepSeek V4.1 Flash has a second, cheaper rate outside the peak window, shown under each one: peak is Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8) and every other hour is off-peak. Which rate a request pays is decided by when it arrives, not when it finishes, so a long generation that crosses the boundary is billed entirely at the rate in force when you sent it. Both rates are in /v1/models as pricing and pricing_offpeak, with the window itself as pricing_window, so a client can price a request before making it.

The discount applies to input and output only. Cache reads are billed at the vendor’s own rate with nothing added and nothing taken off — they already cost between a third and a twenty-fifth of a fresh input token, depending on the model, and discounting them again would save you a rounding error. Nothing in the cache column is struck through for that reason.

Choosing one

  • DeepSeek V4.1 Flash — Fast multimodal model. Reads text, images and video; cheapest per token we sell.
  • DeepSeek V4 Pro — Frontier reasoning model. Best for agents, long-context analysis and code generation.
  • DeepSeek V4 Flash — High-efficiency workhorse. Classification, extraction, chat and high-volume pipelines.
  • GLM 5.3 Flash — Fast multimodal model with enforced structured output. Reads text, images and video.
  • GLM 5.2 — Flagship model for long-horizon agent tasks.

What a request costs

Fresh prompt tokens at the input rate, cached prompt tokens at the cache-read rate, and generated tokens at the output rate. Reasoning tokens are generated tokens and are counted once, inside completion_tokens.

A typical agent turn — a 4,000-token prompt with 2,500 of it cached, producing 600 tokens — is dominated by output, not input. When you are comparing models, compare the output column first.

Reasoning models and max_tokens

DeepSeek V4.1 Flash, DeepSeek V4 Pro, DeepSeek V4 Flash, GLM 5.3 Flash, GLM 5.2 think before answering. Those thinking tokens are generated tokens: they are counted inside completion_tokens, billed at the output rate, and reported separately as reasoning_tokens so you can see the split.

Where the supplier honours it, max_tokens bounds reasoning and answer together: set it too low and the budget is spent thinking before a single word of the answer is produced — you get a 200 response with empty content, billed for the tokens that were generated. Asking DeepSeek V4.1 Flash for a three-word reply with max_tokens: 16 returns exactly that. Where it is not honoured, max_tokens bounds neither: on GLM, max_tokens: 8 returned 784 completion tokens, 775 of them reasoning. So treat it as a bound on the answer rather than as a spending limit — it is also why a thinking request reserves max_tokens plus a measured reasoning allowance. Give a reasoning model room: a few hundred tokens at minimum, and check finish_reason — "length" means it was cut off rather than finished.

If you want short answers without paying for deliberation, pick a model whose Reasoning column below is a dash.

Model ids

Ids are vendor/model, and they are the same strings OpenRouter publishes for the same models — so if you are moving from OpenRouter, the model names in your code do not change either.

"model": "deepseek/deepseek-v4-flash"
"model": "z-ai/glm-5.3-flash"

The un-namespaced ids we published earlier — deepseek-v4-flash, glm-5-3-flash — keep working, and will. So does the namespaced id with the vendor left off, and any capitalisation. GET /v1/models reports the namespaced id as id and the older one as slug; usage rows, key scopes and the URLs on this site are keyed by the slug.

Thinking, and what it does to the bill

Every model in the catalogue reasons before it answers unless you tell it not to, and the reasoning is billed at the output rate. That is the single largest thing a rate card cannot tell you, because it is a difference in how many tokens a task takes rather than what each one costs.

A customer measured it on a JSON extraction: the same six documents cost 2.7 times more on the model with the lower per-token price, because it spent most of its output budget thinking. One document in six never finished — the whole max_tokens budget went on reasoning and the answer came back empty with finish_reason: "length". On a short, mechanical task like extraction, the thinking buys very little.

"reasoning": {"enabled": false}

On the same extraction that one line took a request from 1,047 reasoning tokens to zero — about thirty times cheaper — with identical output. Reach for it whenever the task is mechanical: extraction, classification, reformatting, anything where you already know the shape of the answer. Leave thinking on for the work it is for.

The depth you ask for is passed to the model. "none" and "minimal" ask for no thinking; "low" through "max" leave it on. What the model does with the depth is its own business, and measuring it is harder than it looks — see below.

Median reasoning tokens on a five-person seating puzzle at temperature 0, three runs per value, sending exactly what the gateway sends. Measured 2026-10-03. The prompt matters more than anything else here: a model given a question it can answer without thinking produces almost no reasoning whatever you ask for, which looks exactly like obeying you.

ModelofflowhighmaxOff honoured?
DeepSeek V4 Flash0127331161517869YesOff is off, every run. The depth moves the total but not in a dependable order.
DeepSeek V4.1 Flash0409640964096YesOff is off. On a hard task it reasons to the max_tokens cap at every depth.
DeepSeek V4 Pro010322718211064YesOff is off. The ladder that showed on an easy prompt does not survive a hard one.
GLM 5.20—2973—YesOff is off, every run. Most runs at depth exceeded five minutes and have no figure.
GLM 5.3 Flash401389509—NoReasons whatever you ask. Off produced 134–464 tokens in every spelling; budget for thinking on this model.
Switching thinking off works on every model but one. GLM 5.3 Flash reasons whatever it is asked: on the puzzle above it produced between 134 and 464 reasoning tokens with thinking off, in all three spellings and in the combination this gateway sends. Budget for thinking on that model, and reach for a different one if what you need is a cheap mechanical call.
The depth is not a dependable lever on any of them. An earlier version of this page marked two models as honouring it, measured on an arithmetic question and a three-box riddle. On a problem that actually needs working through, the ordering disappears everywhere — V4 Pro ran 10,322 tokens at "low" and 7,182 at "high". Ask for less thinking if you like; do not plan a bill around getting it.
This table has now been wrong three times, and all three were the same mistake: a measurement taken on one kind of prompt and published as a property of the model. What changed is not the models but the question put to them. The prompt behind these figures is described above so you can disagree with it.
reasoning.max_tokens and reasoning.exclude are still answered with a 400 naming the field, rather than accepted and ignored. A budget we cannot enforce is worse than one we refuse, because you find out about it on the invoice. That is exactly how this was found.

Reasoning tokens are counted in completion_tokens and broken out in completion_tokens_details.reasoning_tokens, so you can see what a task actually spent on thinking before deciding whether to switch it off.

Capabilities

ModelStreamingToolsJSON modeJSON schemaThinks by defaultVision
deepseek/deepseek-v4.1-flash✓✓✓✓✓✓
deepseek/deepseek-v4-pro✓✓✓✓✓—
deepseek/deepseek-v4-flash✓✓✓✓✓—
z-ai/glm-5.3-flash✓✓✓✓✓✓
z-ai/glm-5.2✓✓——✓—

JSON mode is response_format: {"type": "json_object"} — valid JSON, shape up to the model. JSON schema is type: "json_schema", where the shape is yours and the model is held to it. A model without the second tick may accept a schema and ignore it, so the column is measured rather than taken from a datasheet — see structured outputs.

Availability

Every model is served by at least one supplier, and where there is more than one we pick per request on cost and live latency, failing over before the first byte reaches you. A model that no supplier can currently serve returns 503 rather than a silent substitution — we will never answer with a different model from the one you asked for.

Last updated 2026-10-02.