Models and pricing
Prices are per million tokens, billed per token with no rounding up to a minimum. Credit never expires.
The catalogue
| Model id | Input | Output | Cache read | Context |
|---|---|---|---|---|
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash50% off | $0.300$0.150$0.090 off-peak | $1.200$0.600$0.360 off-peak | $0.0060$0.0030 off-peak | 1024K |
DeepSeek V4 Prodeepseek/deepseek-v4-pro50% off | $1.320$0.660 | $3.960$1.980 | $0.044 | 1024K |
DeepSeek V4 Flashdeepseek/deepseek-v4-flash50% off | $0.440$0.220 | $1.320$0.660 | $0.014 | 1024K |
GLM 5.3 Flashz-ai/glm-5.3-flash | $0.150 | $0.500 | $0.030 | 1024K |
GLM 5.2z-ai/glm-5.250% off | $1.400$0.700 | $4.400$2.200 | $0.260 | 1024K |
Struck-through figures are the model vendor’s own published rate; the figure beside them is what we charge. This table is generated from the same catalogue the gateway routes on, so it cannot drift from what you are actually charged.
The figures in the table are the peak rates. DeepSeek V4.1 Flash has a second, cheaper rate outside the peak window, shown under each one: peak is Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8) and every other hour is off-peak. Which rate a request pays is decided by when it arrives, not when it finishes, so a long generation that crosses the boundary is billed entirely at the rate in force when you sent it. Both rates are in /v1/models as pricing and pricing_offpeak, with the window itself as pricing_window, so a client can price a request before making it.
Choosing one
- DeepSeek V4.1 Flash — Fast multimodal model. Reads text, images and video; cheapest per token we sell.
- DeepSeek V4 Pro — Frontier reasoning model. Best for agents, long-context analysis and code generation.
- DeepSeek V4 Flash — High-efficiency workhorse. Classification, extraction, chat and high-volume pipelines.
- GLM 5.3 Flash — Fast multimodal model with enforced structured output. Reads text, images and video.
- GLM 5.2 — Flagship model for long-horizon agent tasks.
What a request costs
Fresh prompt tokens at the input rate, cached prompt tokens at the cache-read rate, and generated tokens at the output rate. Reasoning tokens are generated tokens and are counted once, inside completion_tokens.
Reasoning models and max_tokens
DeepSeek V4.1 Flash, DeepSeek V4 Pro, DeepSeek V4 Flash, GLM 5.3 Flash, GLM 5.2 think before answering. Those thinking tokens are generated tokens: they are counted inside completion_tokens, billed at the output rate, and reported separately as reasoning_tokens so you can see the split.
max_tokens bounds reasoning and answer together: set it too low and the budget is spent thinking before a single word of the answer is produced — you get a 200 response with empty content, billed for the tokens that were generated. Asking DeepSeek V4.1 Flash for a three-word reply with max_tokens: 16 returns exactly that. Where it is not honoured, max_tokens bounds neither: on GLM, max_tokens: 8 returned 784 completion tokens, 775 of them reasoning. So treat it as a bound on the answer rather than as a spending limit — it is also why a thinking request reserves max_tokens plus a measured reasoning allowance. Give a reasoning model room: a few hundred tokens at minimum, and check finish_reason — "length" means it was cut off rather than finished.If you want short answers without paying for deliberation, pick a model whose Reasoning column below is a dash.
Model ids
Ids are vendor/model, and they are the same strings OpenRouter publishes for the same models — so if you are moving from OpenRouter, the model names in your code do not change either.
"model": "deepseek/deepseek-v4-flash"
"model": "z-ai/glm-5.3-flash"The un-namespaced ids we published earlier — deepseek-v4-flash, glm-5-3-flash — keep working, and will. So does the namespaced id with the vendor left off, and any capitalisation. GET /v1/models reports the namespaced id as id and the older one as slug; usage rows, key scopes and the URLs on this site are keyed by the slug.
Thinking, and what it does to the bill
Every model in the catalogue reasons before it answers unless you tell it not to, and the reasoning is billed at the output rate. That is the single largest thing a rate card cannot tell you, because it is a difference in how many tokens a task takes rather than what each one costs.
A customer measured it on a JSON extraction: the same six documents cost 2.7 times more on the model with the lower per-token price, because it spent most of its output budget thinking. One document in six never finished — the whole max_tokens budget went on reasoning and the answer came back empty with finish_reason: "length". On a short, mechanical task like extraction, the thinking buys very little.
"reasoning": {"enabled": false}On the same extraction that one line took a request from 1,047 reasoning tokens to zero — about thirty times cheaper — with identical output. Reach for it whenever the task is mechanical: extraction, classification, reformatting, anything where you already know the shape of the answer. Leave thinking on for the work it is for.
"none" and "minimal" ask for no thinking; "low" through "max" leave it on. What the model does with the depth is its own business, and measuring it is harder than it looks — see below.Median reasoning tokens on a five-person seating puzzle at temperature 0, three runs per value, sending exactly what the gateway sends. Measured 2026-10-03. The prompt matters more than anything else here: a model given a question it can answer without thinking produces almost no reasoning whatever you ask for, which looks exactly like obeying you.
| Model | off | low | high | max | Off honoured? | |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash | 0 | 12733 | 11615 | 17869 | Yes | Off is off, every run. The depth moves the total but not in a dependable order. |
| DeepSeek V4.1 Flash | 0 | 4096 | 4096 | 4096 | Yes | Off is off. On a hard task it reasons to the max_tokens cap at every depth. |
| DeepSeek V4 Pro | 0 | 10322 | 7182 | 11064 | Yes | Off is off. The ladder that showed on an easy prompt does not survive a hard one. |
| GLM 5.2 | 0 | — | 2973 | — | Yes | Off is off, every run. Most runs at depth exceeded five minutes and have no figure. |
| GLM 5.3 Flash | 401 | 389 | 509 | — | No | Reasons whatever you ask. Off produced 134–464 tokens in every spelling; budget for thinking on this model. |
"low" and 7,182 at "high". Ask for less thinking if you like; do not plan a bill around getting it.reasoning.max_tokens and reasoning.exclude are still answered with a 400 naming the field, rather than accepted and ignored. A budget we cannot enforce is worse than one we refuse, because you find out about it on the invoice. That is exactly how this was found.Reasoning tokens are counted in completion_tokens and broken out in completion_tokens_details.reasoning_tokens, so you can see what a task actually spent on thinking before deciding whether to switch it off.
Capabilities
| Model | Streaming | Tools | JSON mode | JSON schema | Thinks by default | Vision |
|---|---|---|---|---|---|---|
deepseek/deepseek-v4.1-flash | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
deepseek/deepseek-v4-pro | ✓ | ✓ | ✓ | ✓ | ✓ | — |
deepseek/deepseek-v4-flash | ✓ | ✓ | ✓ | ✓ | ✓ | — |
z-ai/glm-5.3-flash | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
z-ai/glm-5.2 | ✓ | ✓ | — | — | ✓ | — |
JSON mode is response_format: {"type": "json_object"} — valid JSON, shape up to the model. JSON schema is type: "json_schema", where the shape is yours and the model is held to it. A model without the second tick may accept a schema and ignore it, so the column is measured rather than taken from a datasheet — see structured outputs.
Availability
Every model is served by at least one supplier, and where there is more than one we pick per request on cost and live latency, failing over before the first byte reaches you. A model that no supplier can currently serve returns 503 rather than a silent substitution — we will never answer with a different model from the one you asked for.
Last updated 2026-10-02.