Tokenify

Anthropic Messages

POST /v1/messages. The same models, the same prices, for code that already speaks Anthropic.

Tokenify is not affiliated with Anthropic and does not resell Claude. What we serve here is the Anthropic API format — so a tool written against it can use our models without being rewritten.
from anthropic import Anthropic

client = Anthropic(
    api_key=os.environ["TOKENIFY_API_KEY"],
    base_url="https://api.tokenify.dev",
)

message = client.messages.create(
    model="deepseek/deepseek-v4-flash",
    max_tokens=1024,
    system="You are terse.",
    messages=[{"role": "user", "content": "Why is the sky blue?"}],
)
print(message.content[0].text)

What is supported

FeatureSupportedNotes
System promptYesString or content blocks; becomes the first message upstream.
Multi-turn messagesYesText, image and document blocks.
Tool useYestool_use and tool_result blocks, including parallel calls.
tool_choiceYesauto, any, tool and none all map across.
StreamingYesFull named event sequence — see below.
x-api-keyYesAccepted alongside Authorization: Bearer, so no custom header is needed.
max_tokensRequiredAs in Anthropic’s own API. It also bounds the credit reserved.
thinkingPartly{"type":"enabled"}, {"type":"disabled"} and {"type":"adaptive"} are honoured. thinking.budget_tokens returns 400 — see below.

Claude Code

Claude Code speaks this endpoint and connects with two environment variables:

export ANTHROPIC_BASE_URL=https://api.tokenify.dev
export ANTHROPIC_AUTH_TOKEN=tk-live-…

claude --model 'deepseek/deepseek-v4-flash[1m]'

Tool use, streaming, prompt caching and an exact /context all work. The [1m] suffix matters and there are a few real limits, so it has a page of its own: Claude Code.

Thinking

Thinking is on by default on the models that do it, and thinking: {"type": "disabled"} turns it off for a request. {"type": "adaptive"} — which Claude Code sends on every request — steers nothing and lets the model’s own default stand.

Reasoning comes back as text, not as a thinking block. The models here produce reasoning on a separate channel from the answer, and this endpoint folds it into the same text block rather than inventing a block type we cannot sign. So a response reads reasoning-then-answer in one piece, streamed or not, and content[0].type is "text".

It is billed either way: reasoning tokens are generated tokens and are counted in usage.output_tokens. Until 1 October 2026 a non-streaming reply dropped the reasoning while still charging for it; that was a bug and is fixed, and the two now return the same text for the same request.

thinking.budget_tokens returns 400. A reasoning budget cannot be enforced here — our suppliers accept the field and ignore it — and a cap we cannot apply would bill you for the full reasoning anyway. A customer who sent one was charged thirty times the budget they had written down, for a reply that ran out of room before answering. Send {"type": "disabled"} to switch thinking off instead, or omit thinking to leave it on.

The depth controls on the other dialects behave the same way: models and pricing has what each value does and what it costs.

Streaming events

The full lifecycle, in the order Anthropic defines — blocks are opened, filled and closed one at a time, so an SDK accumulator assembles the message correctly even when the underlying supplier emits parallel tool calls interleaved.

event: message_start
event: content_block_start   (index 0, type "text")
event: content_block_delta   (text_delta)
event: content_block_stop
event: content_block_start   (index 1, type "tool_use")
event: content_block_delta   (input_json_delta)
event: content_block_stop
event: message_delta         (stop_reason + authoritative usage)
event: message_stop

Usage rides on message_delta, including cache_read_input_tokens when the prompt hit a cache.

Errors

Errors come back in the Anthropic envelope, so your client’s retry logic branches on error.type as it expects:

{"type": "error", "error": {"type": "rate_limit_error", "message": "..."}}

A stream that breaks after the headers are out ends with an error event rather than a message_stop, so a truncated answer is never mistaken for a complete one. See errors.

Last updated 2026-10-01.