Anthropic Messages
POST /v1/messages. The same models, the same prices, for code that already speaks Anthropic.
from anthropic import Anthropic
client = Anthropic(
api_key=os.environ["TOKENIFY_API_KEY"],
base_url="https://api.tokenify.dev",
)
message = client.messages.create(
model="deepseek/deepseek-v4-flash",
max_tokens=1024,
system="You are terse.",
messages=[{"role": "user", "content": "Why is the sky blue?"}],
)
print(message.content[0].text)What is supported
| Feature | Supported | Notes |
|---|---|---|
| System prompt | Yes | String or content blocks; becomes the first message upstream. |
| Multi-turn messages | Yes | Text, image and document blocks. |
| Tool use | Yes | tool_use and tool_result blocks, including parallel calls. |
| tool_choice | Yes | auto, any, tool and none all map across. |
| Streaming | Yes | Full named event sequence — see below. |
x-api-key | Yes | Accepted alongside Authorization: Bearer, so no custom header is needed. |
max_tokens | Required | As in Anthropic’s own API. It also bounds the credit reserved. |
thinking | Partly | {"type":"enabled"}, {"type":"disabled"} and {"type":"adaptive"} are honoured. thinking.budget_tokens returns 400 — see below. |
Claude Code
Claude Code speaks this endpoint and connects with two environment variables:
export ANTHROPIC_BASE_URL=https://api.tokenify.dev
export ANTHROPIC_AUTH_TOKEN=tk-live-…
claude --model 'deepseek/deepseek-v4-flash[1m]'Tool use, streaming, prompt caching and an exact /context all work. The [1m] suffix matters and there are a few real limits, so it has a page of its own: Claude Code.
Thinking
Thinking is on by default on the models that do it, and thinking: {"type": "disabled"} turns it off for a request. {"type": "adaptive"} — which Claude Code sends on every request — steers nothing and lets the model’s own default stand.
thinking block. The models here produce reasoning on a separate channel from the answer, and this endpoint folds it into the same text block rather than inventing a block type we cannot sign. So a response reads reasoning-then-answer in one piece, streamed or not, and content[0].type is "text".It is billed either way: reasoning tokens are generated tokens and are counted in usage.output_tokens. Until 1 October 2026 a non-streaming reply dropped the reasoning while still charging for it; that was a bug and is fixed, and the two now return the same text for the same request.
thinking.budget_tokens returns 400. A reasoning budget cannot be enforced here — our suppliers accept the field and ignore it — and a cap we cannot apply would bill you for the full reasoning anyway. A customer who sent one was charged thirty times the budget they had written down, for a reply that ran out of room before answering. Send {"type": "disabled"} to switch thinking off instead, or omit thinking to leave it on.The depth controls on the other dialects behave the same way: models and pricing has what each value does and what it costs.
Streaming events
The full lifecycle, in the order Anthropic defines — blocks are opened, filled and closed one at a time, so an SDK accumulator assembles the message correctly even when the underlying supplier emits parallel tool calls interleaved.
event: message_start
event: content_block_start (index 0, type "text")
event: content_block_delta (text_delta)
event: content_block_stop
event: content_block_start (index 1, type "tool_use")
event: content_block_delta (input_json_delta)
event: content_block_stop
event: message_delta (stop_reason + authoritative usage)
event: message_stopUsage rides on message_delta, including cache_read_input_tokens when the prompt hit a cache.
Errors
Errors come back in the Anthropic envelope, so your client’s retry logic branches on error.type as it expects:
{"type": "error", "error": {"type": "rate_limit_error", "message": "..."}}A stream that breaks after the headers are out ends with an error event rather than a message_stop, so a truncated answer is never mistaken for a complete one. See errors.
Last updated 2026-10-01.