Coming from OpenRouter
The request body, the model ids and the usage accounting are the same. The routing marketplace is not, and that is the difference worth understanding before you compare prices.
The migration
Two lines. The model strings in your code do not change — our ids are the ones OpenRouter publishes for the same models.
- base_url="https://openrouter.ai/api/v1"
- api_key=os.environ["OPENROUTER_API_KEY"]
+ base_url="https://api.tokenify.dev/v1"
+ api_key=os.environ["TOKENIFY_API_KEY"]
# unchanged
model="deepseek/deepseek-v4-flash"provider, models, route, transforms, plugins, preset, usage — the request still succeeds. They are accepted and ignored rather than rejected, so nothing has to be deleted before you try us. What each one does here is spelled out below.What is identical
| Field or behaviour | Notes |
|---|---|
POST /v1/chat/completions | The OpenAI body, forwarded whole. Fields we do not recognise go to the model rather than being dropped. |
Model ids — deepseek/deepseek-v4-flash, z-ai/glm-5.3-flash | The same strings, dots and all. The un-namespaced ids we published earlier also still resolve. |
response_format | Both json_object and json_schema, same shape. Enforced on 4 of our 5 models — see structured outputs. |
supported_parameters | On GET /v1/models, in OpenRouter’s vocabulary, so model-selection logic that filters on it keeps working unchanged. |
usage.cost | In USD, on every response, and on the final frame of a stream. No request field needed — OpenRouter deprecated theirs and we never required one. |
pricing | USD per token as plain decimal strings, with input_cache_read for cache reads. Same field names, same units. |
tools and tool_choice | Standard OpenAI tool calling. Nothing about the shape is ours. |
| Streaming | Server-sent events, usage on the final frame. You do not need stream_options; we ask for it upstream ourselves. |
402 | Out of credit, with the shortfall and the balance in the message. Same status, same meaning. |
response.model | Our own id, echoed back — send it straight into your next request. |
What we do not have
All of these are one feature: OpenRouter is a marketplace that routes between many hosts of the same model, and we are not. Each is accepted in a request and does nothing.
| OpenRouter feature | Here |
|---|---|
provider | No effect. Provider preferences, ordering and allow_fallbacks assume a choice of hosts to express a preference between. |
models and route | No effect. A fallback chain across models is something you can express in your own code; we will not do it silently and bill you for the second attempt. |
transforms | No effect. We do not compress or drop parts of your prompt. What you send is what is billed and what the model sees. |
Variant suffixes — :free, :nitro, :floor, :batch | Not recognised. One endpoint per model at one price, so there is no cheaper or faster variant of it to select. |
reasoning | Partly. enabled is honoured on every model except GLM 5.3 Flash, which reasons whatever it is asked, and every documented effort is accepted: none and minimal mean off, low through max mean on. The depth is forwarded to the model, and what it buys differs by model and by how hard the question is — on a problem that needs working through, the ordering disappears everywhere; the model list has the measurements. The two that ask for a budget we cannot enforce, max_tokens and exclude, return a 400 naming the field. See reasoning. |
| BYOK, and upstream cost in cost_details | Neither. You cannot bring your own upstream key, and we publish what a request cost you rather than what it cost us. |
| The catalogue size | 5 models, chosen and measured, against several hundred there. |
In return: we also serve the Anthropic Messages API over the same models and the same credit, so one key and one balance serve both dialects.
Why the prices look different
This is the part worth reading slowly, because a headline comparison will mislead you in our disfavour and the reason is structural.
A model on OpenRouter is served by fifteen to thirty independent hosts, and the price shown against the model — in their interface, and in the pricing field of their models API — is the cheapest of them. Many of those are quantised: fp8 is eight-bit, fp4 four-bit, against the sixteen the model was trained and published at. Quantisation is a real technique with real uses, and it is also a different product from the one the model vendor ships.
We publish one endpoint per model, at the precision the vendor publishes, at one price. So the comparison that means something is against the hosts on OpenRouter charging the model vendor's own rate, and against those we are competitive — on several models our rate is the same to the cent, and on others we are under it. The comparison against the cheapest quantised host is one we lose, and should: that is not what we are selling.
/api/v1/models/<id>/endpoints; compare against the rows whose quantisation is not fp4 or fp8. Our own rates are on models and pricing, and in GET /v1/models.Three things that move the real bill more than any rate card does. Cache reads are billed at the vendor's published cache rate rather than averaged into a headline number — on a long system prompt that is most of the cost of a conversation. Credit is reserved at the worst case before a request is sent and refunded the moment it settles, so a runaway generation cannot overspend a balance mid-stream. And every model here thinks before answering unless told not to — which GLM 5.3 Flash does anyway, and which on a mechanical task costs thirty times the same task with reasoning: {"enabled": false} — a difference no per-token comparison between us and anybody else will show you.
Which supplier serves a request
We do not publish that, we do not publish how many of them there are, and there is no request parameter to choose one. What we publish instead is what you can act on: the price, the context window and the capabilities of each model, on models and pricing. Who we buy from is our side of the trade; what you are charged and what the model can do is yours.
Every response carries x-tokenify-request-id, and the same value on the body as id. Quote it and we can tell you exactly what happened to that request.
Last updated 2026-10-02.