Coming from Z.AI (GLM)
The model names carry over exactly, which is unusual — you change the base URL and the key. The part worth reading is what thinking does to a GLM bill.
What you change
Two lines. Z.AI's own model names resolve here unchanged, so the strings in your code stay as they are.
- base_url="https://api.z.ai/api/paas/v4"
- api_key=os.environ["ZAI_API_KEY"]
+ base_url="https://api.tokenify.dev"
+ api_key=os.environ["TOKENIFY_API_KEY"]
# unchanged
model="glm-5.3-flash"| Model | Also accepted as |
|---|---|
z-ai/glm-5.3-flash | glm-5.3-flash (Z.AI's own name), glm-5-3-flash, and any capitalisation |
z-ai/glm-5.2 | glm-5.2 (Z.AI's own name), glm-5-2, and any capitalisation |
glm-5.3, glm-4.7, glm-image, cogvideox-3 — is not served here and returns a 404 naming the model, rather than silently routing you somewhere else. The catalogue is the list.Thinking, and what it does to a GLM bill
GLM reasons before answering by default, here as on Z.AI, and the reasoning is billed at the output rate. On GLM the reasoning is also not bounded by your max_tokens: we measured max_tokens: 8 returning 784 completion tokens, 775 of them reasoning. The model's own limits say why — its reasoning ceiling is the same as its output ceiling, so a thinking request can generate up to the model's full output limit whatever you asked for.
Two consequences, and the second will surprise you if nobody says it first.
| What happens | |
|---|---|
| The bill | A short question with a small max_tokens can still cost as if it were a long answer. Switching thinking off took the same request from 784 completion tokens to 9. |
| The credit reserved | Because max_tokens does not bound it, a thinking request is reserved for your max_tokens plus a reasoning allowance — so it can be refused with 402 and a message naming the figure, on a balance that would cover the answer you expected. The allowance is sized from what models actually do, not from the model's documented ceiling, which is its full output limit and would price every thinking request as a maximum-length one. The reservation is released the moment the request settles; only the admission decision is affected. |
"thinking": {"type": "disabled"} // Z.AI's spelling
"reasoning_effort": "none" // also honoured
"reasoning": {"enabled": false} // OpenRouter's spellingAll three work and mean the same thing. On GLM 5.2, thinking off makes max_tokens the bound again and the reservation is priced on it, so small balances stay usable. On GLM 5.3 Flash it does not, because that model reasons anyway (below). Reach for it on mechanical work — extraction, classification, reformatting — and leave thinking on for the work it is for.
reasoning_effort is both translated onto the model's own on/off control and passed through, so "none" and "minimal" switch thinking off and "low" through "max" leave it on at the depth you asked for. The two GLM models differ sharply. GLM 5.2 switches off cleanly — zero reasoning tokens, every run. GLM 5.3 Flash does not switch off at all: measured 2026-10-03 on a seating puzzle that needs working through, it produced 134–464 reasoning tokens with thinking disabled, in all three spellings and in the pair this gateway sends. On a short question it produces almost none, which is why an earlier version of this page called its depth control the strongest in the catalogue — that was the easy prompt talking, not the model. Budget for thinking on 5.3 Flash whatever you send, and note that a request with thinking off there is priced for reasoning for the same reason.Fields with no effect here
Accepted so nothing has to be deleted from a working request, and listed so you do not rely on them. Each was sent to our supplier and measured: it accepts and ignores all of them.
| Z.AI field | Here |
|---|---|
do_sample | No effect. Use temperature and top_p, which are honoured. |
thinking.clear_thinking | No effect. thinking.type is honoured; the sub-field is not. |
request_id | No effect — we assign the id. It comes back as id on the response and as x-tokenify-request-id; quote either at us. |
user_id | No effect. Per-key attribution is in the dashboard instead. |
tool_stream | No effect. Tool calls stream as ordinary deltas. |
finish_reason of sensitive, model_context_window_exceeded, network_error | Not produced. Our supplier has never emitted one across every request we have served — we checked the recorded history rather than assuming. You will see stop, length, content_filter or tool_calls. |
On /v1/messages, a content_filter finish is reported as stop_reason: "refusal", which is the value Anthropic clients branch on.
What is the same
The rest is the OpenAI body you were already sending: messages, temperature, top_p, max_tokens, stop, tools, tool_choice, response_format, stream, and message.reasoning_content with its tokens broken out in completion_tokens_details.reasoning_tokens. Vision takes the same content blocks. And usage.cost is added: what the request cost you, in USD, on every response.
If you are comparing us with OpenRouter as well, our GLM ids are the ones it publishes — z-ai/glm-5.3-flash, z-ai/glm-5.2 — so a switch between the three of us needs no change to the model string. The OpenRouter page covers the rest of that comparison.
Last updated 2026-10-02.