Tokenify

Chat completions

POST /v1/chat/completions. The OpenAI request body, forwarded whole — including fields we have never heard of.

The request

POST https://api.tokenify.dev/v1/chat/completions
Authorization: Bearer $TOKENIFY_API_KEY
Content-Type: application/json

{
  "model": "deepseek/deepseek-v4-flash",
  "messages": [
    {"role": "system", "content": "You are terse."},
    {"role": "user", "content": "Why is the sky blue?"}
  ],
  "max_tokens": 512,
  "temperature": 0.7,
  "stream": false
}

The body is relayed upstream rather than re-serialised from a list of fields we recognise. A sampler flag or tool type added by a supplier next month works here the day they ship it, without a release from us.

Parameters that change the bill

FieldEffect
max_tokensCaps the completion, and caps the credit reserved before the call. Clamped to the model’s own limit.
nOnly 1. Anything above it returns 400: the model returns a single completion whatever is asked for, so reserving credit for more could refuse a request that would otherwise be served. Send the request once per completion you need.
streamDoes not change the price. Usage arrives on the final frame either way.
toolsTool and function schemas are sent to the model and billed as input tokens, like any other part of the prompt.
response_formatA JSON schema is part of the prompt and billed as input tokens, the same as a tool schema. Constraining the output can also shorten it, which lowers the completion charge.
Credit is reserved at the worst case before the request is sent and the unused part is refunded the moment it settles. A single request cannot run away with your balance, and you are never charged for a request a supplier refused. A balance can end up slightly negative — never by more than one request's worth — if a supplier generates more than the max_tokens we asked it for, or if a prompt turns out to be denser than the estimate taken before it was counted.

Tool calls

Standard OpenAI tool calling. Nothing about the shape is ours.

{
  "model": "deepseek/deepseek-v4-flash",
  "messages": [{"role": "user", "content": "What is the weather in Paris?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Current weather for a city",
      "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"]
      }
    }
  }],
  "tool_choice": "auto"
}

JSON and schemas

Two shapes, both through response_format, both identical to OpenAI and OpenRouter.

// valid JSON, shape up to the model
"response_format": {"type": "json_object"}

// valid JSON in exactly this shape
"response_format": {
  "type": "json_schema",
  "json_schema": {
    "name": "person",
    "strict": true,
    "schema": {
      "type": "object",
      "properties": {"full_name": {"type": "string"}, "born": {"type": "integer"}},
      "required": ["full_name", "born"],
      "additionalProperties": false
    }
  }
}

Neither is supported by every model, and a model that does not support one will usually accept the field and ignore it rather than fail — so the list below is generated from the catalogue rather than written by hand, and it is what /v1/models reports.

Enforces a schema (and valid JSON): deepseek/deepseek-v4.1-flash, deepseek/deepseek-v4-pro, deepseek/deepseek-v4-flash, z-ai/glm-5.3-flash.

Neither — these accept response_format and may return prose or fenced JSON anyway: z-ai/glm-5.2. Parse defensively or choose a model above.

Structured outputs covers schemas in full, including how to discover support at runtime.

What a request cost

Every response carries the charge in its usage object, in USD — the same place OpenRouter reports it, so a client that already tracks spend needs no change. On a stream it rides the final frame, the one that carries the token counts.

"usage": {
  "prompt_tokens": 1200,
  "completion_tokens": 400,
  "prompt_tokens_details": {"cached_tokens": 900},
  "cost": 0.000343
}

cost is what came off your balance for this request, to the micro-dollar, after cache reads were priced at the cache rate. It is the figure on your invoice, not an estimate of it. usage: {"include": true} is accepted and unnecessary — OpenRouter deprecated it and we report the cost either way.

The response also names the model as model, using our id rather than the model vendor's internal snapshot name. You can send it straight back in your next request.

Reasoning models

Models that think before answering return their reasoning in message.reasoning_content and count it in completion_tokens_details.reasoning_tokens. Those tokens are already inside completion_tokens, so they are billed at the output rate exactly once.

Every model here thinks by default, and on a mechanical task that can be most of the bill — thirty times most, measured. Switch it off with the same field OpenRouter uses:

"reasoning": {"enabled": false}
FieldEffect
reasoning: {"enabled": false}No thinking. Honoured on every model but one: GLM 5.3 Flash reasons anyway.
reasoning: {"enabled": true}Thinking on, which is also what happens if you send nothing.
reasoning: {"effort": "low" | "medium" | "high" | "max"}Thinking on, at the depth you asked for: the word is passed to the model. How much it changes is the model's to decide and differs between them — models and pricing has what each one was measured to do.
reasoning: {"effort": "none" | "minimal"}Thinking off, the same as enabled: false.
reasoning: {"max_tokens": n}400. The reasoning budget cannot be capped apart from max_tokens.
reasoning: {"exclude": true}400. Hiding the reasoning would not reduce what it costs.
The three that return 400 are all accepted and silently ignored by some gateways. They were by us too, until a customer sent max_tokens: 64, was billed for two thousand reasoning tokens, and got an empty answer back. GET /v1/models lists reasoning under supported_parameters for the models that honour it.

Listing models

curl https://api.tokenify.dev/v1/models -H "Authorization: Bearer $TOKENIFY_API_KEY"

Returns the models this key may call, in the OpenAI listing shape. See models and pricing for what each one costs.

Last updated 2026-10-02.