Anthropic-compatible
Keep the Claude SDK. Change the base URL. Pay a fraction.
Tokenify serves the Anthropic Messages API at /v1/messages — same request shape, same streaming events, same error envelope. Your Claude code runs unchanged against DeepSeek V4.
| Per 1M tokens | Input | Output |
|---|---|---|
| Claude Haiku 4.5 | $1.000 | $5.000 |
| DeepSeek V4 Pro | $0.660 | $1.980 |
| DeepSeek V4.1 Flash | $0.150 | $0.600 |
| DeepSeek V4.1 Flash off-peak | $0.090 | $0.360 |
Claude Haiku 4.5 price from claude.com, checked 2026-10-01. Tokenify serves DeepSeek models, not Claude.
One line
The same code, a different base URL
Messages, system prompts, content blocks, tool use, streaming and errors are all translated. Nothing in your client changes.
from anthropic import Anthropic
client = Anthropic(
base_url="https://api.tokenify.dev", # the only line that changes
api_key=os.environ["TOKENIFY_API_KEY"],
)
message = client.messages.create(
model="deepseek/deepseek-v4-pro",
max_tokens=1024,
system="Be concise.",
messages=[{"role": "user", "content": "Explain B-trees"}],
)
print(message.content[0].text)What it costs
Roughly 8× cheaper per blended token
Blended at a 3:1 prompt-to-completion ratio. Anthropic's published list price is shown for comparison; we do not sell it.
| Model | Input / 1M | Output / 1M | Cache read / 1M | Context | |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash deepseek/deepseek-v4.1-flash50% off | $0.300$0.150 $0.090 off-peak | $1.200$0.600 $0.360 off-peak | $0.006 $0.003 off-peak | 1049K | Use this model |
| DeepSeek V4 Pro deepseek/deepseek-v4-pro50% off | $1.320$0.660 | $3.960$1.980 | $0.044 | 1049K | Use this model |
| DeepSeek V4 Flash deepseek/deepseek-v4-flash50% off | $0.440$0.220 | $1.320$0.660 | $0.014 | 1049K | Use this model |
| GLM 5.3 Flash z-ai/glm-5.3-flash | $0.150 | $0.500 | $0.030 | 1049K | Use this model |
| GLM 5.2 z-ai/glm-5.250% off | $1.400$0.700 | $4.400$2.200 | $0.260 | 1049K | Use this model |
- Input / 1M
- $0.300$0.150$0.090 off-peak
- Output / 1M
- $1.200$0.600$0.360 off-peak
- Cache read / 1M
- $0.006$0.003 off-peak
- Context
- 1049K
- Input / 1M
- $1.320$0.660
- Output / 1M
- $3.960$1.980
- Cache read / 1M
- $0.044
- Context
- 1049K
- Input / 1M
- $0.440$0.220
- Output / 1M
- $1.320$0.660
- Cache read / 1M
- $0.014
- Context
- 1049K
- Input / 1M
- $0.150
- Output / 1M
- $0.500
- Cache read / 1M
- $0.030
- Context
- 1049K
- Input / 1M
- $1.400$0.700
- Output / 1M
- $4.400$2.200
- Cache read / 1M
- $0.260
- Context
- 1049K
Struck-through figures are the model vendor’s own published rate; the price beside them is what we charge. Billed per token, no minimum. Cache reads are the exception and carry no discount — they already cost between a third and a twenty-fifth of the input rate, and we pass the supplier’s cache rate through rather than flattening it into one headline number.
Off-peak pricing. DeepSeek V4.1 Flash is billed at two rates. Peak hours are Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8); everything else is off-peak, including weekends. A request is billed at the rate in force when it is made, and every request in your logs records which one it was.
For reference, Claude Haiku 4.5 lists at $1.00 per million input tokens and $5.00 output, a blended $2.00 — from claude.com, checked 2026-10-01. DeepSeek V4.1 Flash through Tokenify blends to $0.262. They are different models with different capabilities — the point is that switching between them costs you one line, so you can measure the trade on your own workload instead of guessing at it.
What is and is not supported
Translated
- system as a top-level field
- Text and image content blocks
- tool_use and tool_result round trips
- tool_choice, including auto, any and a named tool
- stop_sequences, temperature, top_p
- Streaming lifecycle events, including tool JSON deltas
- x-api-key auth and anthropic-version
- Anthropic-shaped error envelopes
Not available
- Anthropic’s own models — we do not resell them
- Anthropic-specific beta headers and features
- The Batches and Files APIs
- Prompt caching control headers (cache hits are still billed at the cache rate when a supplier reports them)
Last updated 2026-10-01. Prices come from the live catalogue at the time this page was built.