Quickstart
Two minutes from here to your first completion. If you already call OpenAI or Anthropic, one line changes.
1. Get a key
Create an account and one is issued during setup. Keys are shown once, so put it straight into your environment:
export TOKENIFY_API_KEY="tk-live-..."2. Add some credit
Requests are billed per token against prepaid credit, and a request with no credit behind it is refused with a 402. New accounts get a small amount to try the API with; after that, top up by card or stablecoin. Credit does not expire.
3. Make a request
The base URL is https://api.tokenify.dev/v1 and the key goes in an Authorization: Bearer header. Nothing else about your code changes.
from openai import OpenAI
client = OpenAI(
api_key=os.environ["TOKENIFY_API_KEY"],
base_url="https://api.tokenify.dev/v1", # the only line that changes
)
response = client.chat.completions.create(
model="deepseek/deepseek-v4-flash",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)4. Read the response
It is the shape your SDK already expects, with two additions worth knowing about. Cached prompt tokens are reported separately, because they are billed at a fraction of fresh ones, and two headers come back: the canonical model id we resolved your request to, and a request id.
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "deepseek/deepseek-v4-flash",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello! How can I help?" },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 9,
"completion_tokens": 8,
"total_tokens": 17,
"prompt_tokens_details": { "cached_tokens": 0 },
"completion_tokens_details": { "reasoning_tokens": 0 }
}
}x-tokenify-model: deepseek/deepseek-v4-flash
x-tokenify-request-id: req_aba350ea35840912b70756beQuote x-tokenify-request-id in a support conversation and we can find the exact request, including what it cost and how it was served. The same id is on the response body as id, prefixed chatcmpl-, so one string covers your logs and ours.
What to read next
- Streaming — tokens as they are generated, and where the usage figures arrive.
- Prompt caching — how a long system prompt gets much cheaper on the second turn.
- Models and pricing — what each one costs and what it can do.
Last updated 2026-09-28.