Tokenify documentation
Frontier models at wholesale rates, behind the API your code already speaks. Keep your SDK; change the base URL and the key.
Tokenify is an inference gateway. It exposes the OpenAI Chat Completions API at /v1/chat/completions and the Anthropic Messages API at /v1/messages, routes each request to whichever supplier can serve it cheapest and fastest, and bills you per token against prepaid credit. There is no subscription and no minimum.
curl https://api.tokenify.dev/v1/chat/completions \
-H "Authorization: Bearer $TOKENIFY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4-flash",
"messages": [{"role": "user", "content": "Hello"}]
}'If you already call OpenAI, that is the whole migration: point base_url at https://api.tokenify.dev/v1, use a Tokenify key, and name a Tokenify model. Everything else — streaming, tools, JSON mode, the response shape — is unchanged.
Start here
For agents and crawlers
Every page here is static HTML with no client-side rendering, so a plain fetch returns the same words a browser shows. If you are pointing a coding agent at this API, these two files are written for it:
- /llms.txt — the base URL, the auth header, the model ids and the endpoints, in about forty lines.
- /llms-full.txt — the same, plus every parameter, error code and worked example, as one plain-text document.
You will need a key to make a request. Creating an account takes a minute and issues one immediately.
Last updated 2026-09-28.