Streaming
Set stream: true and read server-sent events. Nothing between you and the supplier buffers them.
The frames
data: {"id":"chatcmpl-1","choices":[{"delta":{"role":"assistant"},"index":0}]}
data: {"id":"chatcmpl-1","choices":[{"delta":{"content":"Hello"},"index":0}]}
data: {"id":"chatcmpl-1","choices":[{"delta":{},"finish_reason":"stop","index":0}]}
data: {"id":"chatcmpl-1","choices":[],"usage":{"prompt_tokens":9,"completion_tokens":8,
"prompt_tokens_details":{"cached_tokens":0}}}
data: [DONE]Usage arrives on the last frame
You do not need to ask for it. OpenAI requires stream_options: {"include_usage": true} to get token counts on a stream; we set it on every streaming request whether you send it or not, because we bill on the supplier’s reported counts and will not bill on an estimate while a real number is available for the asking.
Nothing buffers
The edge is configured to flush every frame as it arrives, with no response buffering and no write deadline, and it sends X-Accel-Buffering: no so an intermediate proxy does not hold your tokens hostage either. A reasoning model that thinks for two minutes before its first token will not be cut off.
Cancelling
Close the connection. That is an ordinary thing to do — a user pressing stop, a mobile app going to the background, an SDK timeout — and it is not counted against the supplier’s health, so aborting generations will never push your traffic onto a more expensive route.
You are billed for what the supplier generated up to that point, because they generated it and they charge us for it. Cancelling early genuinely costs less; it does not cost nothing.
# Python: stop reading and close
with client.chat.completions.create(
model="deepseek/deepseek-v4-flash",
messages=[{"role": "user", "content": "Write an essay"}],
stream=True,
) as stream:
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
if enough():
break # closing the stream stops the charge growingFailover and the first byte
If a supplier fails before the first byte reaches you, the request is retried on the next one and you never see it. Once the first byte is out, failover is no longer possible — the response has started — so a mid-stream supplier failure surfaces as a truncated answer with an error, not as a silent switch.
On the Anthropic endpoint that surfaces as an error event rather than a clean message_stop, so an SDK can tell a finished answer from an interrupted one.
curl
curl -N https://api.tokenify.dev/v1/chat/completions \
-H "Authorization: Bearer $TOKENIFY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"deepseek/deepseek-v4-flash","stream":true,
"messages":[{"role":"user","content":"Count to five"}]}'-N matters: without it curl buffers the response and you will see the whole answer at once no matter what we do.
Last updated 2026-09-28.