Tokenify

Coming from DeepSeek

The request body is the same one, the thinking control is the same field, and vision takes the same content blocks. The model names are ours, and two beta features are not here.

What you change

The base URL and the key, and the model name. We accept the base URL with or without /v1, so whichever shape your client has works.

- base_url="https://api.deepseek.com"
- api_key=os.environ["DEEPSEEK_API_KEY"]
+ base_url="https://api.tokenify.dev"
+ api_key=os.environ["TOKENIFY_API_KEY"]

- model="deepseek-flash"
+ model="deepseek/deepseek-v4-flash"
The model name is the one thing that cannot pass through. DeepSeek serves its own models under its own names, and we serve several vendors — so ids here are namespaced, and deepseek-chat, deepseek-reasoner and deepseek-flash resolve to nothing. GET /models lists what does.
If you were usingUse
deepseek-v4-prodeepseek/deepseek-v4-pro
deepseek-flashdeepseek/deepseek-v4-flash, or deepseek/deepseek-v4.1-flash if you need vision
deepseek-chat / deepseek-reasonerOlder names with no equivalent here — see the catalogue and pick by what you need

The un-namespaced ids we published earlier — deepseek-v4-pro, deepseek-v4-flash — also still resolve, which is why model="deepseek-v4-pro" happens to work unchanged from DeepSeek.

What works unchanged

Field or behaviourNotes
POST /chat/completionsWith or without /v1. Same body, forwarded whole — fields we do not recognise reach the model rather than being dropped.
thinkingDeepSeek's own control. {"type": "disabled"} and {"type": "enabled"} are honoured. thinking.budget_tokens returns a 400 — see below.
reasoning_effort"none" and "minimal" switch thinking off; "low", "medium", "high" and "max" leave it on at the depth you asked for — the word is passed to the model as well as being translated onto its on/off control. What the depth buys differs by model, and on this family only V4 Pro honours it: six runs on a logic puzzle gave 102 reasoning tokens at "low" rising to 343 at "max", ordered and with the ranges apart. On V4 Flash and V4.1 Flash the rungs are noise — "max" measured below "low" on one prompt. Switching thinking off works on all three, every time. Measured 2026-10-03; the model list has the figures. Gated on the same per-model measurement as reasoning and thinking.
response_formatBoth json_object and json_schema. Which models enforce a schema is measured per model, not assumed.
tools, tool_choiceStandard OpenAI tool calling.
logprobs, top_logprobsForwarded and returned.
Vision — the same content blockstext and image_url blocks, base64 data URLs and remote https URLs, and the detail field. Images in a system message are a 400, as on DeepSeek.
message.reasoning_contentReturned when the model thinks, and its tokens are broken out in completion_tokens_details.reasoning_tokens.
StreamingServer-sent events, usage on the final frame, terminated by data: [DONE]. You do not need stream_options; we ask for usage upstream ourselves.
frequency_penalty, presence_penaltyAccepted. DeepSeek documents them as ignored; we forward them and make no claim either way.

What is different

DeepSeekHere
Model namesNamespaced, and ours. 3 DeepSeek models plus other vendors — see the catalogue.
thinking.budget_tokens400, naming the field. A depth can be asked for and is passed on, but a token budget is not enforced upstream — measured, it is accepted and then overrun — and a budget we cannot enforce would bill you for the full reasoning anyway. Send {"type": "disabled"} instead.
The /beta base URL — chat prefix completion, prefix: true, reasoning_content as inputNot served. There is no /beta here and those fields have no effect.
FIM completionNot served.
system_fingerprintNot returned.
Request body up to 48 MiB16 MiB. Large vision payloads may need resizing or fewer images per request.
idOurs, as chatcmpl-<request id>, and the same value is on the x-tokenify-request-id header. Quote it at us and we can find the request.
usage.costExtra, not missing: every response says in USD what the request cost you.

Thinking is on by default here too

Same default as DeepSeek: every model reasons before answering unless told not to, and the reasoning is billed at the output rate. Telling it not to works on all three DeepSeek models here, every run; it is one of the Z.AI models that reasons regardless, which models and pricing records. On mechanical work — extraction, classification, reformatting — that is most of the bill, and it can consume the whole max_tokens budget and return empty content with finish_reason: "length".

"thinking": {"type": "disabled"}

Either spelling works — DeepSeek's thinking or OpenRouter's reasoning: {"enabled": false} — so you do not have to change it when you arrive. Models and pricing has the measured numbers.

Last updated 2026-10-02.