# Tokenify > Wholesale LLM inference behind an OpenAI-compatible API, billed per token > against prepaid credit. > Cache reads carry no discount: they are billed at the model vendor's published > cache rate rather than averaged into a headline number. ## Integrating Base URL: https://api.tokenify.dev/v1 Auth header: Authorization: Bearer Keys: created at https://www.tokenify.dev/app/keys — shown once, prefixed tk-live- or tk-test- Drop-in for the OpenAI SDK: set base_url and api_key, change nothing else. Also serves the Anthropic Messages API at https://api.tokenify.dev/v1/messages, which additionally accepts the x-api-key header. ## Endpoints - POST https://api.tokenify.dev/v1/chat/completions — OpenAI Chat Completions, streaming supported - POST https://api.tokenify.dev/v1/messages — Anthropic Messages, streaming supported. Reasoning is folded into the text block rather than returned as a `thinking` block, in both the streamed and non-streamed reply, and is billed inside `usage.output_tokens` - POST https://api.tokenify.dev/v1/responses — OpenAI Responses, streaming supported. Send `store: false` and the full `input` each turn: conversations are not retained, so `previous_response_id` is refused. This is the endpoint Codex speaks. - POST https://api.tokenify.dev/v1/messages/count_tokens — exact prompt size, free and unbilled, counted with the same tokenizer the invoice uses - GET https://api.tokenify.dev/v1/models — models this key may call ## Models - deepseek/deepseek-v4.1-flash — in $0.150/M (list $0.300/M), out $0.600/M (list $1.200/M), cache read $0.0060/M (not discounted), 1024K context, 50% off list on input and output; OFF-PEAK in $0.090/M, out $0.360/M, cache read $0.0030/M (peak is Mon-Fri 09:00-12:00 and 14:00-18:00 UTC+8; billed at the rate when the request is made), reasoning model, enforces JSON schema, json_object needs "json" in the prompt - deepseek/deepseek-v4-pro — in $0.660/M (list $1.320/M), out $1.980/M (list $3.960/M), cache read $0.044/M (not discounted), 1024K context, 50% off list on input and output, reasoning model, enforces JSON schema - deepseek/deepseek-v4-flash — in $0.220/M (list $0.440/M), out $0.660/M (list $1.320/M), cache read $0.014/M (not discounted), 1024K context, 50% off list on input and output, reasoning model, enforces JSON schema - z-ai/glm-5.3-flash — in $0.150/M, out $0.500/M, cache read $0.030/M (not discounted), 1024K context, reasoning model, enforces JSON schema - z-ai/glm-5.2 — in $0.700/M (list $1.400/M), out $2.200/M (list $4.400/M), cache read $0.260/M (not discounted), 1024K context, 50% off list on input and output, reasoning model ## Model ids Ids are vendor/model — deepseek/deepseek-v4-flash, z-ai/glm-5.3-flash — and are the same strings OpenRouter publishes for the same models, so model names need no change when switching. The earlier un-namespaced ids (deepseek-v4-flash, glm-5-3-flash) still resolve, as does the namespaced id with the vendor omitted, in any capitalisation. GET /v1/models reports the namespaced id as `id`. Prices in that response are USD per token, as plain decimal strings with no exponent: `"prompt": "0.00000015"` is $0.15 per million tokens. The cache-read price is `input_cache_read`. Some models are billed at two rates by the clock. For those, `pricing` is the peak rate and `pricing_offpeak` carries the cheaper one, with `pricing_window` giving the hours and a `peak_now` flag for the moment you asked. A request is billed at the rate in force when it is made, so work that can wait is cheaper run outside the window. Models without `pricing_offpeak` have one rate at all hours. ## What a request cost Every response carries the charge in `usage.cost`, in USD — what actually came off the balance for that request, cache reads priced at the cache rate. On a stream it is on the final frame, the one carrying the token counts. The `usage: {"include": true}` request field is accepted and unnecessary; the cost comes back either way. `response.model` is our own model id, not the model vendor's internal snapshot name, so it can be sent straight back in the next request. ## Structured outputs Models marked "enforces JSON schema" accept `response_format` with `type: "json_schema"` and return output conforming to the schema. Models without that mark may accept the field and ignore it, returning whatever shape the prompt suggested rather than an error — validate the result, or choose a marked model. A model marked "schema support not measured" is one this file could not get an answer for; ask `GET /v1/models`, which is always current. `GET /v1/models` carries the same fact as `supported_parameters`, which contains `structured_outputs` for exactly the models that enforce a schema and `response_format` for those plus the models that return valid JSON without enforcing a shape. A model that accepts the field and ignores it lists neither. Every model accepts `response_format` — the request body is forwarded whole — so treat the list as what is obeyed, not what is accepted. ## Reasoning models Models marked "reasoning model" above think before answering. Those tokens are counted inside completion_tokens and billed at the output rate; the split is reported as usage.completion_tokens_details.reasoning_tokens. Where it is honoured upstream, max_tokens bounds reasoning and answer together: set it too low and the whole budget is spent thinking, so the response is a 200 with empty content, billed for the tokens generated, and finish_reason is "length". Where it does not, max_tokens bounds neither — on GLM, measured, max_tokens: 8 returned 784 completion tokens, 775 of them reasoning. Treat max_tokens as a bound on the answer, not as a spending limit; the reservation we take is max_tokens plus a measured reasoning allowance for that reason — on any request that leaves thinking on, and on every request to a model measured to reason with thinking off. Give a reasoning model a few hundred output tokens at minimum, and check finish_reason before treating an empty string as a real answer. ## Thinking costs output tokens Every model reasons before answering unless told otherwise, and the reasoning is billed at the output rate. Telling it otherwise works on every model but one — see the table below. On a mechanical task — extraction, classification, reformatting — that can be most of the bill, and it can consume the whole `max_tokens` budget and return empty content with `finish_reason: "length"`. Measured on one extraction: 1,047 reasoning tokens against 30 with thinking off, ~30x the cost, identical output. Switch it off per request: "reasoning": {"enabled": false} Honoured on every model but one: GLM 5.3 Flash reasons anyway, and the table below has the figures. `reasoning.effort` of "none" or "minimal" means off, and "low", "medium", "high" or "max" means on at that depth — the word you send is passed to the model, not just translated into on/off. These return 400 with the field named, rather than being ignored: `reasoning.max_tokens` and `reasoning.exclude` — a budget that cannot be enforced is worse than one that is refused. Measured 2026-10-03: three runs per value on a five-person seating puzzle at temperature 0, sending exactly what this gateway sends. Median reasoning tokens. The prompt matters more than anything else here — a model given a question it can answer without thinking produces almost no reasoning whatever you ask for, which looks exactly like obeying you. model off low high max off honoured deepseek/deepseek-v4-flash 0 12733 11615 17869 yes deepseek/deepseek-v4.1-flash 0 4096 4096 4096 yes deepseek/deepseek-v4-pro 0 10322 7182 11064 yes z-ai/glm-5.2 0 — 2973 — yes z-ai/glm-5.3-flash 401 389 509 — no deepseek/deepseek-v4-flash: Off is off, every run. The depth moves the total but not in a dependable order. deepseek/deepseek-v4.1-flash: Off is off. On a hard task it reasons to the max_tokens cap at every depth. deepseek/deepseek-v4-pro: Off is off. The ladder that showed on an easy prompt does not survive a hard one. z-ai/glm-5.2: Off is off, every run. Most runs at depth exceeded five minutes and have no figure. z-ai/glm-5.3-flash: Reasons whatever you ask. Off produced 134–464 tokens in every spelling; budget for thinking on this model. Switching thinking off works on every model but one. GLM 5.3 Flash reasons whatever it is asked: 134-464 reasoning tokens with thinking off, in all three spellings and in the pair this gateway sends. Budget for thinking there, or pick another model for a cheap mechanical call. The depth is not a dependable lever on any of them. Asking for less is accepted and forwarded; on a problem that actually needs working through, the ordering disappears — V4 Pro ran 10,322 tokens at "low" and 7,182 at "high". Ask for less if you like; do not plan a bill around getting it. Reasoning tokens are inside `completion_tokens` and broken out in `completion_tokens_details.reasoning_tokens`. ## The model does not see response_format `response_format` constrains decoding. It is not added to the prompt and the model never sees the schema — same here, on OpenAI and on OpenRouter. A prompt that refers to "the schema" without describing it is therefore ambiguous to the model, and on a thinking model that ambiguity is billed. Measured, 14 runs, same schema, max_tokens 2000: - "Read this into the schema: ..." — 564 to 2000 reasoning tokens, median 1326, 1 run in 14 spent the whole budget reasoning and returned empty content with finish_reason "length". - "Extract full_name, born and field ... return only JSON matching that shape: ..." — 62 to 116 reasoning tokens, median 76, no failures. So name the fields in the prompt, and send `"reasoning": {"enabled": false}` for mechanical work. The prompt is never rewritten for you: doing so would change the token count and the result without being asked. Models that additionally require the word "json" in the prompt when using `response_format: {"type":"json_object"}` are marked in the model list; it does not apply to `json_schema`. ## Coming from OpenRouter Compatible: POST /v1/chat/completions with the OpenAI body; the same model ids, dots included; `response_format` with `json_object` and `json_schema`; `supported_parameters` on GET /v1/models in the same vocabulary; `usage.cost` in USD on every response; `pricing` as USD per token with `input_cache_read`; `tools`/`tool_choice`; SSE streaming with usage on the final frame; 402 when credit runs out. Changing the base URL and the key is the whole migration. Accepted and ignored, not rejected, so nothing needs deleting first: `provider`, `models`, `route`, `transforms`, `plugins`, `preset`, `usage`. Variant suffixes (:free, :nitro, :floor, :batch) are not recognised. `reasoning` is honoured for on/off — see "Thinking costs output tokens" below for what is honoured and what returns 400. On price: a model on OpenRouter is served by 15-30 independent hosts and the price shown against the model is the cheapest of them, many quantised to fp8 or fp4. We publish one endpoint per model at the precision the vendor publishes, at one price. Compare against the hosts charging the vendor's own rate — OpenRouter lists every host and its quantisation at /api/v1/models//endpoints — not against the cheapest quantised row. We do not publish which supplier serves a request, we do not publish how many of them there are, and there is no parameter to choose one. What is published is the price, the context window and the capabilities of every model, at https://www.tokenify.dev/docs/models. Full comparison: https://www.tokenify.dev/docs/openrouter/ ## Coming from DeepSeek Change the base URL and the key; `/v1` is optional, both shapes work. The request body carries over unchanged, including `thinking`, `response_format`, `tools`, `logprobs`, and vision content blocks (base64 data URLs, remote https URLs, `detail`). Images in a system message are a 400, as on DeepSeek. `message.reasoning_content` and `completion_tokens_details.reasoning_tokens` are returned as there. Model names do NOT carry over: `deepseek-chat`, `deepseek-reasoner` and `deepseek-flash` resolve to nothing here. Use the namespaced ids from the model list above — `deepseek/deepseek-v4-flash`, `deepseek/deepseek-v4.1-flash` for vision, `deepseek/deepseek-v4-pro`. `reasoning_effort` is honoured: "none" and "minimal" disable thinking, "low"/"medium"/"high"/"max" leave it on and are passed to the model as the depth. Switching thinking off works on all three of these, every run: zero reasoning tokens, in every spelling, including on a problem that needs working through. The depth is a different matter — on that problem the ordering disappears, with V4 Pro running 10,322 tokens at "low" and 7,182 at "high". Ask for less if you like; do not plan a bill around getting it. All three spellings — `reasoning`, `thinking`, `reasoning_effort` — follow the same rule, and `reasoning` is translated onto `reasoning_effort` because that is the only spelling the models actually read. Not served: the `/beta` base URL and everything on it (chat prefix completion, `prefix: true`, `reasoning_content` as input), FIM completion, and `system_fingerprint`. `thinking.budget_tokens` returns 400 — a depth is passed on but a budget is not enforced upstream, so send `thinking: {"type":"disabled"}` or an effort instead. `n` above 1 returns 400: the supplier returns one completion whatever is asked for, so reserving credit for more could refuse a request that would otherwise be served. Request bodies are capped at 16 MiB rather than 48 MiB. Extra: `usage.cost` in USD on every response. Full comparison: https://www.tokenify.dev/docs/deepseek/ ## Coming from Z.AI (GLM) Change the base URL and the key. Z.AI's own GLM model names resolve unchanged; so do OpenRouter's `z-ai/...` ids. Models we do not carry (glm-5.3, glm-4.7, glm-image, cogvideox-3) return 404 naming the model. On GLM, `max_tokens` does NOT bound a thinking request: measured, `max_tokens: 8` returned 784 completion tokens, 775 of them reasoning, because the model vendor's reasoning ceiling equals its output ceiling. Two consequences: the bill can be far above what max_tokens suggests, and a thinking request is reserved for max_tokens plus a reasoning allowance — so it can return 402 naming a figure on a balance that would cover the answer you expected. On GLM 5.2, switching thinking off makes max_tokens the bound again and the reservation shrinks with it. On GLM 5.3 Flash it does not, because that model reasons anyway (below), so the allowance is reserved there whatever you send: "thinking": {"type": "disabled"} // Z.AI's spelling "reasoning_effort": "none" // also honoured "reasoning": {"enabled": false} // OpenRouter's spelling All three mean the same thing here. `reasoning_effort` is both translated onto the on/off control and passed through as the depth. The two GLM models differ, and the difference is the whole point of the table above. GLM 5.2 switches off cleanly: zero reasoning tokens, every run. **GLM 5.3 Flash does not switch off at all** — on a problem that needs working through it produced 134-464 reasoning tokens with thinking disabled, in all three spellings and in the pair this gateway sends. On a short question it produces almost none, which is why an earlier version of this file called it the strongest depth control in the catalogue; that was the easy prompt talking, not the model. Budget for thinking on 5.3 Flash whatever you send. OpenRouter's `reasoning` object is translated onto `reasoning_effort` rather than forwarded, because the models do not read it: on GLM 5.3 Flash, `reasoning: {"effort":"low"}` and `{"effort":"max"}` both produce that model's default. Accepted and ignored (measured upstream): `do_sample`, `thinking.clear_thinking`, `request_id`, `user_id`, `tool_stream`. Z.AI's `finish_reason` values `sensitive`, `model_context_window_exceeded` and `network_error` are never produced here; expect stop, length, content_filter or tool_calls. On /v1/messages a content_filter finish is reported as `stop_reason: "refusal"`. Full comparison: https://www.tokenify.dev/docs/zai/ ## Things worth knowing - Usage is always reported on streams; you do not need stream_options. - Cached prompt tokens appear in usage.prompt_tokens_details.cached_tokens and are billed at the cache-read rate, which is a small fraction of input. - Credit is reserved before a call and the unused part refunded on settlement, so a runaway generation cannot empty a balance. One exception: reasoning is not bounded by max_tokens, so a thinking request reserves max_tokens plus a measured reasoning allowance and an unusually long reasoning run can settle slightly above what was reserved, taking the balance slightly negative. Send `reasoning: {"enabled": false}` (or `thinking: {"type":"disabled"}`, or `reasoning_effort: "none"`) and max_tokens bounds the whole request again on every model that honours it — not on GLM 5.3 Flash, which reasons anyway and is reserved for it whatever you send. - 402 insufficient_credit means top up; 429 carries Retry-After; 502 upstream_unavailable was not charged. - Log the x-tokenify-request-id response header: it is what identifies a request in support. ## Documentation - https://www.tokenify.dev/docs/quickstart/ — first request, four languages - https://www.tokenify.dev/docs/authentication/ — keys, scopes, rotation - https://www.tokenify.dev/docs/chat-completions/ — parameters, tools, JSON mode - https://www.tokenify.dev/docs/streaming/ — SSE frames, usage, cancellation - https://www.tokenify.dev/docs/anthropic/ — the Messages API - https://www.tokenify.dev/docs/prompt-caching/ — how to earn cache hits - https://www.tokenify.dev/docs/errors/ — every status and whether to retry - https://www.tokenify.dev/docs/limits/ — rate limits and quotas - https://www.tokenify.dev/docs/codex/ — point OpenAI Codex here (Responses API) - https://www.tokenify.dev/docs/claude-code/ — point Claude Code here (Messages API) - https://www.tokenify.dev/claude-code/ — the Claude Code setup guide: the variables, what was measured to work, what a session costs - https://www.tokenify.dev/docs/models/ — every model, both pricing windows, capabilities - https://www.tokenify.dev/pricing/ — the rate card as a page, with how billing works - https://www.tokenify.dev/llms-full.txt — everything above plus worked examples, as one file