Structured outputs
Send a JSON schema and get back JSON shaped like it. The same response_format OpenAI and OpenRouter use, with a per-model answer about whether it is enforced or merely encouraged.
Sending a schema
Pass response_format with type: "json_schema". The shape is the one OpenAI defined and OpenRouter follows, so code written against either works here unchanged.
from openai import OpenAI
client = OpenAI(base_url="https://api.tokenify.dev/v1", api_key="tk-live-…")
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Ada Lovelace, born 1815, mathematician."}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "person",
"strict": True,
"schema": {
"type": "object",
"properties": {
"full_name": {"type": "string"},
"born": {"type": "integer"},
"fields": {"type": "array", "items": {"type": "string"}},
},
"required": ["full_name", "born", "fields"],
"additionalProperties": False,
},
},
},
# Extraction does not need the model to think about it, and thinking is
# billed at the output rate. See /docs/models#reasoning.
extra_body={"reasoning": {"enabled": False}},
)
print(resp.choices[0].message.content)
# {"full_name":"Augusta Ada King","born":1815,"fields":["mathematics"]}response_format: {"type": "json_object"} also works, and asks only for valid JSON without constraining its shape.
Which models enforce it
Not every model that accepts a schema obeys one. Some treat it as a strong hint and return whatever shape the prompt suggested — no error, just the wrong data. That is the worst failure mode there is, because your code has nothing to catch, so we test each model against a schema that deliberately contradicts its prompt and publish the result rather than the datasheet.
Schema enforced: deepseek/deepseek-v4.1-flash, deepseek/deepseek-v4-pro, deepseek/deepseek-v4-flash, z-ai/glm-5.3-flash.
Not enforced — these accept response_format and may ignore it: z-ai/glm-5.2. Validate what comes back, or pick a model from the list above.
json_object wants the word “json”
With response_format: {"type": "json_object"}, deepseek/deepseek-v4.1-flash rejects the request unless the prompt contains the word json somewhere — 400, with a message saying so. No other model in the catalogue enforces it, and OpenAI documents the same rule for its own json_object, so code written against OpenAI is usually already fine. Code written against a model that does not enforce it is not, and finds out on the first request after switching.
It does not apply to json_schema, which is the mode to prefer anyway.
The model never sees your schema
This is the one thing on this page that will cost you money if you skip it.
response_format constrains how the answer is decoded. It is not added to your prompt, and the model is never shown it — here, on OpenAI, or on OpenRouter. So a prompt like "read this into the schema" is, from the model's side, a request referring to something it cannot see.
On a model that thinks, that ambiguity is expensive. We measured fourteen runs of exactly that prompt: the model spent between 564 and 2,000 reasoning tokens, median 1,326, working out what "the schema" might mean — its own reasoning reads "need infer task… maybe schema.org… could be a table?" — before the decoder forced valid JSON out of it anyway. Adding the field names to the prompt took the same fourteen runs to a median of 76 reasoning tokens. Seventeen times less, for one sentence.
| Prompt | Reasoning tokens (14 runs) | Failures |
|---|---|---|
"Read this into the schema: …" | 564 – 2,000, median 1,326 | 1 in 14 |
"Extract full_name, born and field… return only JSON matching that shape: …" | 62 – 116, median 76 | none |
The failures are the part that looks like flakiness rather than cost. When the guessing runs past max_tokens, the budget is gone before a single character of answer is produced: you get finish_reason: "length", empty content, and a bill for the full completion. Because the median sits well under the ceiling and the tail crosses it, it fails occasionally rather than consistently — and the runs that succeed look perfect, so nothing points at the cause.
"reasoning": {"enabled": false}, which takes reasoning to zero and the same batch to about a thirtieth of the cost. The example above does both.We do not fix this for you by quietly appending the schema to your prompt. It would change your token count and your results without your asking, and we do not rewrite prompts. Full detail on thinking and its cost is on models and pricing.
Asking the API instead
You do not have to read this page to find out. GET /v1/models carries supported_parameters for every model, in the same vocabulary OpenRouter uses, so a client that already filters on it needs no changes.
curl -s https://api.tokenify.dev/v1/models \
| jq '.data[] | select(.supported_parameters | index("structured_outputs")) | .id'structured_outputs appears only for models measured to enforce a schema. response_format appears for those and for models measured to return valid JSON without enforcing a shape. A model that accepts the field and ignores it lists neither — every model in the catalogue accepts response_format, because we forward the body whole, so a list built from what is accepted would say nothing at all. The distinction is the point.
Coming from OpenRouter
supported_parameters means your model-selection logic keeps working. Coming from OpenRouter covers the rest of the comparison, including what we do not have and why a headline price comparison misleads.One difference worth knowing: OpenRouter routes a single model id across several providers, so it offers provider.require_parameters to avoid landing on an endpoint that lacks a feature. Here /v1/models is the contract: what it says a model supports is what you get on every request, so there is no routing decision left for you to constrain.
Chat completions covers the rest of the request body.
Last updated 2026-09-28.