Chat completions
POST /v1/chat/completions. The OpenAI request body, forwarded whole — including fields we have never heard of.
The request
POST https://api.tokenify.dev/v1/chat/completions
Authorization: Bearer $TOKENIFY_API_KEY
Content-Type: application/json
{
"model": "deepseek/deepseek-v4-flash",
"messages": [
{"role": "system", "content": "You are terse."},
{"role": "user", "content": "Why is the sky blue?"}
],
"max_tokens": 512,
"temperature": 0.7,
"stream": false
}The body is relayed upstream rather than re-serialised from a list of fields we recognise. A sampler flag or tool type added by a supplier next month works here the day they ship it, without a release from us.
Parameters that change the bill
| Field | Effect |
|---|---|
max_tokens | Caps the completion, and caps the credit reserved before the call. Clamped to the model’s own limit. |
n | Only 1. Anything above it returns 400: the model returns a single completion whatever is asked for, so reserving credit for more could refuse a request that would otherwise be served. Send the request once per completion you need. |
stream | Does not change the price. Usage arrives on the final frame either way. |
tools | Tool and function schemas are sent to the model and billed as input tokens, like any other part of the prompt. |
response_format | A JSON schema is part of the prompt and billed as input tokens, the same as a tool schema. Constraining the output can also shorten it, which lowers the completion charge. |
max_tokens we asked it for, or if a prompt turns out to be denser than the estimate taken before it was counted.Tool calls
Standard OpenAI tool calling. Nothing about the shape is ours.
{
"model": "deepseek/deepseek-v4-flash",
"messages": [{"role": "user", "content": "What is the weather in Paris?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}],
"tool_choice": "auto"
}JSON and schemas
Two shapes, both through response_format, both identical to OpenAI and OpenRouter.
// valid JSON, shape up to the model
"response_format": {"type": "json_object"}
// valid JSON in exactly this shape
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "person",
"strict": true,
"schema": {
"type": "object",
"properties": {"full_name": {"type": "string"}, "born": {"type": "integer"}},
"required": ["full_name", "born"],
"additionalProperties": false
}
}
}Neither is supported by every model, and a model that does not support one will usually accept the field and ignore it rather than fail — so the list below is generated from the catalogue rather than written by hand, and it is what /v1/models reports.
Enforces a schema (and valid JSON): deepseek/deepseek-v4.1-flash, deepseek/deepseek-v4-pro, deepseek/deepseek-v4-flash, z-ai/glm-5.3-flash.
Neither — these accept response_format and may return prose or fenced JSON anyway: z-ai/glm-5.2. Parse defensively or choose a model above.
Structured outputs covers schemas in full, including how to discover support at runtime.
What a request cost
Every response carries the charge in its usage object, in USD — the same place OpenRouter reports it, so a client that already tracks spend needs no change. On a stream it rides the final frame, the one that carries the token counts.
"usage": {
"prompt_tokens": 1200,
"completion_tokens": 400,
"prompt_tokens_details": {"cached_tokens": 900},
"cost": 0.000343
}cost is what came off your balance for this request, to the micro-dollar, after cache reads were priced at the cache rate. It is the figure on your invoice, not an estimate of it. usage: {"include": true} is accepted and unnecessary — OpenRouter deprecated it and we report the cost either way.
model, using our id rather than the model vendor's internal snapshot name. You can send it straight back in your next request.Reasoning models
Models that think before answering return their reasoning in message.reasoning_content and count it in completion_tokens_details.reasoning_tokens. Those tokens are already inside completion_tokens, so they are billed at the output rate exactly once.
Every model here thinks by default, and on a mechanical task that can be most of the bill — thirty times most, measured. Switch it off with the same field OpenRouter uses:
"reasoning": {"enabled": false}| Field | Effect |
|---|---|
reasoning: {"enabled": false} | No thinking. Honoured on every model but one: GLM 5.3 Flash reasons anyway. |
reasoning: {"enabled": true} | Thinking on, which is also what happens if you send nothing. |
reasoning: {"effort": "low" | "medium" | "high" | "max"} | Thinking on, at the depth you asked for: the word is passed to the model. How much it changes is the model's to decide and differs between them — models and pricing has what each one was measured to do. |
reasoning: {"effort": "none" | "minimal"} | Thinking off, the same as enabled: false. |
reasoning: {"max_tokens": n} | 400. The reasoning budget cannot be capped apart from max_tokens. |
reasoning: {"exclude": true} | 400. Hiding the reasoning would not reduce what it costs. |
400 are all accepted and silently ignored by some gateways. They were by us too, until a customer sent max_tokens: 64, was billed for two thousand reasoning tokens, and got an empty answer back. GET /v1/models lists reasoning under supported_parameters for the models that honour it.Listing models
curl https://api.tokenify.dev/v1/models -H "Authorization: Bearer $TOKENIFY_API_KEY"Returns the models this key may call, in the OpenAI listing shape. See models and pricing for what each one costs.
Last updated 2026-10-02.