Claude Code
Use Claude Code with DeepSeek — no router needed
Claude Code speaks the Anthropic Messages API, and so do we. Export three environment variables and it runs against DeepSeek V4 Pro with no proxy in between, no second process to keep alive and the whole agent loop intact — tools, streaming and caching included.
export ANTHROPIC_BASE_URL=https://api.tokenify.dev export ANTHROPIC_AUTH_TOKEN=tk-live-… export ANTHROPIC_MODEL='deepseek/deepseek-v4-pro[1m]'
Setup
Three variables, then run claude
Nothing is installed and nothing runs in the background. Claude Code already knows how to talk to this API; it only needs to be told where it is.
- 1
Create an API key
Sign up, then generate a key in the dashboard. No waitlist and no regional card checks.
- 2
Export three variables
ANTHROPIC_BASE_URL points Claude Code at this API, ANTHROPIC_AUTH_TOKEN authenticates, and ANTHROPIC_MODEL names the model to use. A fourth, ANTHROPIC_DEFAULT_HAIKU_MODEL, sets the cheaper model used for small background tasks.
- 3
Keep the [1m] suffix on the model name
Claude Code reads it to learn the real context window and strips it before sending the request. Without it the conversation is compacted at 200K on a model that holds a million tokens.
- 4
Check it answers
Run claude -p "say ok" in any directory. A reply means the key, the base URL and the model name are all correct.
export ANTHROPIC_BASE_URL=https://api.tokenify.dev
export ANTHROPIC_AUTH_TOKEN=tk-live-…
export ANTHROPIC_MODEL='deepseek/deepseek-v4-pro[1m]'
export ANTHROPIC_DEFAULT_HAIKU_MODEL='deepseek/deepseek-v4-flash[1m]'
claudeclaude-sonnet-4-5 comes back as Unknown model rather than being quietly answered by something else. That is why ANTHROPIC_MODEL is not optional. The [1m] suffix is read by Claude Code and stripped before the request reaches us, so it changes nothing about routing or price — it only stops the conversation being compacted at 200K on a model that holds a million tokens.Which models
One model for the work, a cheaper one for the errands
Claude Code uses a second, smaller model for background tasks such as titling a conversation. Both are worth setting.
| Variable | Model | Input | Output | Cache read |
|---|---|---|---|---|
ANTHROPIC_MODELthe one that does the work | DeepSeek V4 Prodeepseek/deepseek-v4-pro | $0.660 | $1.980 | $0.044 |
ANTHROPIC_DEFAULT_HAIKU_MODELsmall background tasks | DeepSeek V4 Flashdeepseek/deepseek-v4-flash | $0.220 | $0.660 | $0.014 |
Per 1M tokens. Both models hold 1M tokens of context, which is why the [1m] suffix matters. Every price on this page comes from the live catalogue — see pricing for the full list.
Measured
What works, and what does not
Every row below was run against this API with the real client. Nothing here is inferred from the protocol.
| Feature | Status | What we found |
|---|---|---|
| Conversation and streaming | Verified | A real session over three turns, streamed token by token, with keep-alives during long pauses. |
| Tool use | Verified | Read, Edit, Glob, Grep, Bash and the rest. Two Read calls issued in parallel came back correctly, and an Edit rewrote the file on disk. Tool names survive every turn. |
| Thinking with tools, over several turns | Verified | No 400 and no lost reasoning. Claude Code sends thinking: {"type":"adaptive"} and the model decides; an explicit {"type":"enabled"} works too, through a tool_use and tool_result round trip. |
| Prompt caching | Verified | Automatic — no cache_control to send. The measured session read 16,128 tokens from cache, 95% of its prompt, billed at the cache rate and itemised per request. |
| Exact /context | Verified | We serve /v1/messages/count_tokens against the model’s own tokenizer, so the gauge matches the invoice. |
| metadata.user_id | Verified | Accepted and ignored. It does not change routing or billing, and it does not error. |
| Subagents and /compact | Verified | Ordinary requests as far as this API is concerned. |
| Cache-write accounting | Limit | Cache reads are reported; the write that populates the cache is not broken out separately and bills as ordinary input. Your total is right, the breakdown has one fewer line than Anthropic’s. |
| The /model picker | Limit | It lists only model ids containing “claude” or “anthropic”, so ours never appear. Name the model in ANTHROPIC_MODEL instead. |
| The “unrecognized model” warning | Limit | Printed at startup because our ids are not in the client’s catalogue. Cosmetic; the [1m] suffix handles the part that matters. |
| Reasoning depth | Limit | The depth is passed to the model on the OpenAI dialect, and this client never sends one — it asks for adaptive thinking with no budget, so the model’s own default stands. Switching thinking off still works. |
| Images | Limit | Works on DeepSeek V4.1 Flash and GLM 5.3 Flash, which read images on both the Anthropic and OpenAI formats — a screenshot pasted into Claude Code is answered correctly. The other models are text only, so point ANTHROPIC_MODEL at one of these if you paste images. |
| Server-side web search | No | Not available on this API yet. Claude Code’s own WebSearch and WebFetch tools are unaffected — those run on your machine. |
| Anthropic beta capabilities | No | Context management, the effort beta and the structured-output betas are features of Anthropic’s API rather than of the models we serve. Dropped quietly, so a request carrying them still runs. |
Checked 2026-10-01 with Claude Code 2.1.285, against DeepSeek V4 Pro except the image row, which was measured on DeepSeek V4.1 Flash and GLM 5.3 Flash. The environment variable names change between releases of the client, so if yours is much newer and something here is wrong, the reference page is where we keep the current spelling.
Compared
With and without a router
claude-code-router and tools like it exist to translate between the Anthropic API and a provider that does not speak it. We speak it, so the translation layer has nothing to do.
Straight at this API
- Three environment variables and no install
- No proxy process to start, supervise or restart
- One place for a bug to be, not two
- Usage is counted in the Anthropic shape you sent, so the dashboard matches /cost
- Streaming and tool calls are relayed, not reassembled
What a router still gives you
- Several providers at once, switchable inside one session
- Per-request rules — a different model for background tasks than for plans
- Providers that do not offer an Anthropic-compatible endpoint at all
- Rewriting requests locally before they leave your machine
If you already run a router and like it, nothing here asks you to stop — point it at https://api.tokenify.dev as one more Anthropic-format provider. The claim on this page is narrower than “routers are bad”: if the only reason yours exists is to reach a cheaper model from Claude Code, you can delete it.
What it costs
A measured session, not an estimate
Agent traffic is unusual: it resends the system prompt, the tool schemas and every file it has read on each turn, so most of the prompt is a cache read rather than fresh input.
3 turns on DeepSeek V4 Pro: read two files with the Read tool, then answer. What the gateway recorded:
- Prompt tokens
- 16,911
- Of those, cached
- 16,128 (95%)
- Output tokens
- 275
- Charged
- $0.0018
Measured 2026-10-01; priced from the catalogue above. Your sessions will be longer than this one, and the shape is what carries over rather than the total: the cache rate is what an agent mostly pays, which is why a cheap cache read matters more here than a cheap input token.
The same workload is costed across every seller of these models on the DeepSeek pricing comparison — its “coding agent” scenario is this shape at a month’s volume.
Credit is reserved before each request against the model’s output ceiling and the unused part refunded the moment it finishes, so the reservation is a floor on your balance rather than a cost. An agent pointed at a nearly empty balance will stop mid-task with a 402 that says how much it needed.
Questions
Frequently asked
Does this use Anthropic’s Claude models?
No. Claude Code is a client that speaks the Anthropic Messages API, and that is the API format we serve. The model answering you is DeepSeek or GLM, whichever you name in ANTHROPIC_MODEL. Tokenify is independent and does not resell Anthropic’s models.
Why do I have to set ANTHROPIC_MODEL?
Because we do not map Claude model names onto other vendors’ models. A request for claude-sonnet-4-5 returns “Unknown model” rather than quietly answering with something else, so you have to name the model you want. Renaming our models to look like Anthropic’s would make the model picker work and would be a lie about whose models they are.
Do I still need claude-code-router?
No. A router exists to translate between the Anthropic API and something that does not speak it. We speak it natively, so there is no proxy to run, no extra process to keep alive and no second place for a bug to live. A router is still the right tool if you want to switch between several providers from inside one Claude Code session.
Does prompt caching work?
Yes, and it is most of what an agent costs. Caching is automatic — you do not send cache_control — and cache reads are billed at the cache rate and itemised per request. On the session measured for this page, the prompt came back 95% cached.
Why does Claude Code say “unrecognized model” when it starts?
Its built-in catalogue lists Anthropic’s models and ours are not in it. The warning is cosmetic and the session works. The part that does matter — the context window it assumes — is what the [1m] suffix on the model name is for.
What happens when my balance runs out mid-task?
The request returns 402 and says how much it needed. Credit is reserved before each request and the unused part refunded when it finishes, so an agent pointed at a nearly empty balance stops rather than overdrawing.
Writing against the Anthropic SDK rather than running Claude Code? The Claude API page covers that, and the Messages API docs have the endpoint in full.
Last updated 2026-10-01. Prices come from the live catalogue at the time this page was built.