Tokenify

DeepSeek V4.1 Flash API

Fast multimodal model. Reads text, images and video; cheapest per token we sell.

$0.150 per 1M input and $0.600 per 1M output at peak, and $0.090 / $0.360 off-peak. 1M context. 50% below DeepSeek's published rate.

Call it in three lines

from openai import OpenAI

client = OpenAI(
    base_url="https://api.tokenify.dev/v1",
    api_key=os.environ["TOKENIFY_API_KEY"],
)

response = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

What it costs

A typical agent turn — a 4,000-token prompt with 2,500 of it cached, producing 600 tokens — costs $0.00060. A million of those turns costs $600.

Input / 1M
$0.300$0.150
Output / 1M
$1.200$0.600
Cache read / 1M
$0.006

50% offoff DeepSeek’s published rate on input and output. Cache reads pass through undiscounted — they already cost a fraction of a fresh token.

Off-peak pricing

Peak hours are Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8). Everything else — nights, weekends and the two hours over lunch — is off-peak, and a request is billed at the rate in force when it is made. Your request log records which rate applied.

Input / 1M
$0.150$0.090
Output / 1M
$0.600$0.360
Cache read / 1M
$0.003

How it compares

ModelInputOutputCache readContext
DeepSeek V4.1 Flash$0.150$0.600$0.0061M
DeepSeek V4 Pro$0.660$1.980$0.0441M
DeepSeek V4 Flash$0.220$0.660$0.0141M

Peak rates, per 1M tokens, from the live catalogue. All models · the full rate card · how sellers of this model compare.

Against the other fast models

ModelInputOutputCache readAPI format
GPT-6 LunaOpenAI$0.100$0.500$0.010OpenAI
Claude Haiku 4.5Anthropic$1.000$5.000$0.100Anthropic
Gemini 3.8 FlashGoogle$0.750$3.750$0.075Google
DeepSeek V4.1 FlashTokenify, peak$0.150$0.600$0.006OpenAI and Anthropic
DeepSeek V4.1 FlashTokenify, off-peak$0.090$0.360$0.003OpenAI and Anthropic

Per 1M tokens. Our rows come from the live catalogue at build time. Other vendors’ prices were read from their own pages and last checked 2026-10-01: openai.com (2026-09-30), claude.com (2026-10-01), ai.google.dev (2026-10-01).

Gemini 3.8 Flash is on a promotional rate until 2026-12-31; it becomes $1.50 / $7.50 / $0.15 on 1 January 2027.

Questions about this model

Is this the same model as DeepSeek's own API?

Yes — DeepSeek V4.1 Flash is DeepSeek's model, served through our endpoint rather than theirs. We do not fine-tune, quantise or substitute it. What differs is the price, the API formats on offer and how it is billed.

What are the peak hours?

Peak is Mon-Fri 09:00-12:00 and 14:00-18:00 (UTC+8). Every other hour — nights, weekends and the two hours over lunch — is off-peak, and a request is billed at the rate in force when it is made. Your request log records which rate applied.

How much cheaper is this than DeepSeek's list price?

50% on input and output, blended three parts prompt to one part completion. Cache reads are not discounted: they are billed at DeepSeek's published cache rate, which is already a small fraction of a fresh token.

Can I use it with Claude Code?

Yes. DeepSeek V4.1 Flash answers on the Anthropic Messages API as well as the OpenAI one, so Claude Code reaches it with three environment variables and no router. The Claude Code page has the setup and the compatibility results.

Do I need a subscription?

No. Billing is per token from prepaid credit, with no subscription, no seats and no minimum. Credit does not expire.

Last updated 2026-10-01. Prices come from the live catalogue at the time this page was built.