Tokenify
New5 models live

Wholesale inference,
through the SDK you already use.

DeepSeek and Zhipu models from $0.150 per million input tokens. One endpoint that speaks both the OpenAI and the Anthropic API formats. Pay per token, no subscription, automatic failover across suppliers.

No card required · Credit never expires

< 3ms
routing overhead
2
API formats
99.9%
uptime target
90 days
request history

Why Tokenify

Cheaper is the easy part

Anyone can discount a token. Doing it without adding latency, losing requests or making your bill unauditable is the part that takes building.

You keep your SDK

Tokenify speaks both the OpenAI and the Anthropic wire formats. Change one base URL in code you have already written and shipped — no client library to swap, no response shape to re-parse, no migration branch.

Wholesale pricing, per token

We buy capacity in volume and resell it per token. There is no subscription, no seat, no minimum and no monthly commitment. Credit never expires, and what you do not spend is refundable.

Cache discounts you actually receive

Cached prompt tokens are billed at a fraction of fresh ones, itemised separately on every invoice line. Most resellers flatten that into one headline rate and keep the difference. On agent workloads with long stable system prompts it is often the largest saving of all.

Failover before the first byte

Sources are ranked per request on live cost and latency. One that starts failing is pulled within seconds and the request retried, before a single byte has reached you.

Billing you can audit

Every request records its token counts, the exact rate applied — including which pricing window it fell in — and what you were charged, kept for 90 days. Requests we had to estimate because a supplier reported no usage are flagged as estimated rather than quietly mixed in.

Spending that cannot run away

Credit is reserved before a request and reconciled after it, so a runaway generation stops at your balance instead of past it. Each API key carries its own rate limits, so one bad batch job cannot starve production.

How it works

Three steps, and the second one is a single line

There is no SDK to install, no proxy to run and no migration to plan.

  1. 1

    Create a key

    Sign up with email or Google and generate an API key from the dashboard. It is shown once — store it in your environment, not in source control.

  2. 2

    Change the base URL

    Point your existing OpenAI or Anthropic client at Tokenify and pass the key. Everything else in your code stays exactly as it is.

  3. 3

    Watch it land

    Requests, tokens, latency and the exact charge appear in your dashboard within a minute. Top up when you want to; nothing renews on its own.

Quickstart

Your code already knows how to talk to us

Tokenify speaks the OpenAI wire format. Point your existing client at a different base URL and keep everything else.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.tokenify.dev/v1",   # the only line that changes
    api_key=os.environ["TOKENIFY_API_KEY"],
)

stream = client.chat.completions.create(
    model="deepseek/deepseek-v4-pro",
    messages=[{"role": "user", "content": "Explain B-trees briefly."}],
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

What passes straight through

  • Two API dialects
    OpenAI at /v1/chat/completions, Anthropic at /v1/messages.
  • Server-sent event streaming
    Relayed byte for byte, flushed per chunk.
  • Tool and function calling
    Including parallel calls and strict schemas.
  • JSON and structured output
    response_format is forwarded untouched.
  • Every sampler parameter
    temperature, top_p, stop, seed, penalties.
  • Error envelopes
    Same shape your SDK already parses.
  • Response headers
    Plus the canonical model, your balance and a request id we add.

Catalog

5 models today, more as we sign suppliers

One endpoint for all of them, with a failing source out of rotation within seconds and the retry made before any bytes reach your client.

DeepSeek V4.1 Flash

deepseek/deepseek-v4.1-flash
50% off

Fast multimodal model. Reads text, images and video; cheapest per token we sell.

Input
$0.300$0.150/M
$0.090 off-peak
Output
$1.200$0.600/M
$0.360 off-peak
Cache read
$0.006/M
$0.003 off-peak

Peak is Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8); every other hour is off-peak, priced by when the request arrives.

DeepSeek V4.1 Flash API details →

DeepSeek V4 Pro

deepseek/deepseek-v4-pro
50% off

Frontier reasoning model. Best for agents, long-context analysis and code generation.

Input
$1.320$0.660/M
Output
$3.960$1.980/M
Cache read
$0.044/M
DeepSeek V4 Pro API details →

DeepSeek V4 Flash

deepseek/deepseek-v4-flash
50% off

High-efficiency workhorse. Classification, extraction, chat and high-volume pipelines.

Input
$0.440$0.220/M
Output
$1.320$0.660/M
Cache read
$0.014/M
DeepSeek V4 Flash API details →

GLM 5.3 Flash

z-ai/glm-5.3-flash

Fast multimodal model with enforced structured output. Reads text, images and video.

Input
$0.150/M
Output
$0.500/M
Cache read
$0.030/M
GLM 5.3 Flash API details →

GLM 5.2

z-ai/glm-5.2
50% off

Flagship model for long-horizon agent tasks.

Input
$1.400$0.700/M
Output
$4.400$2.200/M
Cache read
$0.260/M
GLM 5.2 API details →

Pricing

Per token. Nothing else.

No subscription, no seats, no minimum. Add credit, spend it, add more when you want to.

DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash
50% off
Input / 1M
$0.300$0.150
$0.090 off-peak
Output / 1M
$1.200$0.600
$0.360 off-peak
Cache read / 1M
$0.006
$0.003 off-peak
Context
1049K
Use this model
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
50% off
Input / 1M
$1.320$0.660
Output / 1M
$3.960$1.980
Cache read / 1M
$0.044
Context
1049K
Use this model
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
50% off
Input / 1M
$0.440$0.220
Output / 1M
$1.320$0.660
Cache read / 1M
$0.014
Context
1049K
Use this model
GLM 5.3 Flash
z-ai/glm-5.3-flash
Input / 1M
$0.150
Output / 1M
$0.500
Cache read / 1M
$0.030
Context
1049K
Use this model
GLM 5.2
z-ai/glm-5.2
50% off
Input / 1M
$1.400$0.700
Output / 1M
$4.400$2.200
Cache read / 1M
$0.260
Context
1049K
Use this model

Struck-through figures are the model vendor’s own published rate; the price beside them is what we charge. Billed per token, no minimum. Cache reads are the exception and carry no discount — they already cost between a third and a twenty-fifth of the input rate, and we pass the supplier’s cache rate through rather than flattening it into one headline number.

Off-peak pricing. DeepSeek V4.1 Flash is billed at two rates. Peak hours are Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8); everything else is off-peak, including weekends. A request is billed at the rate in force when it is made, and every request in your logs records which one it was.

How that compares

Blended at a 3:1 prompt-to-completion ratio, which is roughly what production traffic looks like.

ModelInputOutputCache readAPI format
GPT-6 LunaOpenAI$0.100$0.500$0.010OpenAI
Claude Haiku 4.5Anthropic$1.000$5.000$0.100Anthropic
Gemini 3.8 FlashGoogle$0.750$3.750$0.075Google
DeepSeek V4.1 FlashTokenify, peak$0.150$0.600$0.006OpenAI and Anthropic
DeepSeek V4.1 FlashTokenify, off-peak$0.090$0.360$0.003OpenAI and Anthropic

Per 1M tokens. Our rows come from the live catalogue at build time. Other vendors’ prices were read from their own pages and last checked 2026-10-01: openai.com (2026-09-30), claude.com (2026-10-01), ai.google.dev (2026-10-01).

Gemini 3.8 Flash is on a promotional rate until 2026-12-31; it becomes $1.50 / $7.50 / $0.15 on 1 January 2027.

No tiers, no minimum, no expiry

The listed discount applies to the first token and the hundred-millionth alike. Credit never expires and unspent credit is refundable.

Platform

Built to sit in front of production traffic

Reselling tokens is easy. Doing it without adding latency, losing money or dropping requests is the hard part.

Sub-3ms routing overhead

Authentication, rate limiting and credit reservation resolve in a single atomic Redis operation. Your time-to-first-token is the supplier’s, plus a memcpy.

Failover by default

Every request is ranked across the sources we hold for that model, by price and by live latency. A circuit breaker pulls a failing one within seconds and puts it back once it recovers.

Per-request cost visibility

Every response carries a request id and the canonical model it resolved to. That id turns into token counts, the rate applied and what you were charged, per request, for 90 days.

Cache priced at the vendor rate

Cached prompt tokens bill at a fraction of the input rate, at the model vendor’s published cache rate rather than averaged into a headline number. On agent workloads with long stable system prompts this is often the largest single line of savings.

Keys with their own limits

Issue a key per service with its own RPM and TPM ceiling. Revocation takes effect across the fleet in seconds, not at the next cache expiry.

Credit that cannot overdraw

We reserve the worst-case cost before the call and refund the difference when it lands. A runaway generation stops at your balance instead of past it.

FAQ

Questions worth answering before you sign up

Is Tokenify really a drop-in replacement for the OpenAI SDK?

Yes. Change base_url to https://api.tokenify.dev/v1 and swap the API key. Streaming, tool calling, JSON mode and every sampler parameter pass through unchanged, and errors come back in the same envelope your client already parses.

How can the prices be lower than DeepSeek list price?

We buy inference capacity in volume at wholesale rates and resell it per token. Our routing layer sends each request to the cheapest healthy source that serves the model, so the saving is structural rather than a launch promotion.

What happens if a supplier goes down?

A circuit breaker takes the failing source out of rotation within seconds and the request is retried before any bytes reach your client, so where a model has another source you see a slightly slower response instead of an error. Where every source has been tried and none answered, you get a 502 — or a 429 if that is what the last attempt said — naming the model, and the request is not charged.

How is billing calculated?

Per token, against the exact usage the supplier reports. Cache-hit prompt tokens are billed at the cache-read rate, not the input rate. Credit is reserved when a request starts and the unused portion is refunded the moment it finishes, so a runaway generation cannot empty your balance. Reasoning models are the one exception worth knowing: their reasoning is not bounded by max_tokens, so the reservation includes a measured allowance for it and an unusually long run can settle slightly above what was held. Send reasoning: {"enabled": false} on mechanical work and, on every model that honours it, max_tokens bounds the whole request again — GLM 5.3 Flash reasons anyway and is reserved for it whatever you send.

Is there a subscription or minimum spend?

No. You add credit, the credit never expires, and you spend it per token. There is no monthly fee, no seat pricing and no commitment.

Do you store my prompts?

We store metadata for your dashboard — model, token counts, latency, cost and status — for 90 days. Prompt and completion content is streamed through and never written to disk.

What rate limits apply?

Each API key carries its own requests-per-minute and tokens-per-minute limits, which you set in the dashboard. That means one runaway batch job cannot starve your production traffic. Higher limits are available on request.

How do I pay?

Card through Stripe, or USDT/USDC and other major cryptocurrencies. Auto-reload keeps production running when the balance runs low, and credit never expires.

Change one line. Keep the rest of your stack.

Create an account, get free credit, and send your first request in under five minutes.

Get your API key