Codex and the Responses API
POST /v1/responses. The endpoint Codex speaks, and the config that points it here.
Codex talks to its model provider over the Responses API and, since codex-cli 0.154.0, nothing else: wire_api = "chat" is refused at startup. So pointing it at a provider means pointing it at a Responses endpoint, which is what /v1/responses is.
Add a provider block to ~/.codex/config.toml and select it. The API key comes from an environment variable, so it never lands in the file:
[model_providers.tokenify]
name = "Tokenify"
base_url = "https://api.tokenify.dev/v1"
wire_api = "responses"
env_key = "TOKENIFY_API_KEY"
model_provider = "tokenify"
model = "deepseek/deepseek-v4-flash"export TOKENIFY_API_KEY=tk-live-…
codex "explain what this repo does"/v1/models, which serves the OpenAI shape that every other client expects. It is a warning, not a failure, and everything works underneath it.What the endpoint supports
| Feature | Supported | Notes |
|---|---|---|
| Streaming | ✓ | The full typed event set, relayed as the model produces it. |
| Function tools | ✓ | Arguments stream incrementally, as they do on OpenAI. |
| Reasoning summaries | ✓ | reasoning: { summary: "auto" } — asking to see the thinking, which every model here allows. |
| Reasoning effort | Per model | reasoning.effort is the control, and is accepted only by models listing reasoning under supported_parameters. See models and pricing. |
| Non-streaming | ✓ | Send stream: false and get the response object in one piece. |
| store / previous_response_id | — | Refused. Conversations are not retained; send the full input each turn. |
| Built-in tools | — | web_search, file_search, code_interpreter and the rest are removed from the request rather than passed on. |
Conversations are not stored
previous_response_id returns a 400, and store is always false. We do not retain your conversations, which means there is nothing on our side for a response id to point at — and a response id is the only thing that would authorise reading one back.
In practice this costs nothing. A Responses client that cannot store sends the whole conversation on every turn, which is exactly what Codex does by default. If you are writing your own client, keep the input array yourself and append to it.
Web search is not available
Codex includes web_search in the tool list of every request. We serve models, not a search index, and the model vendor’s own web search is a server-side tool our account is not yet entitled to use — so those entries are removed from the request before it goes on.
They have to be removed rather than passed along: a request carrying a built-in tool the vendor will not run is rejected in its entirety, so forwarding it would fail every Codex request rather than merely leaving search unavailable. Your own function tools are untouched and work normally, which is how Codex does its actual work — running commands, reading and editing files.
The practical effect is that Codex believes web search is available and the model never calls it. Asked to search, it says so rather than inventing an answer. We would rather ship this properly than soon: it is billed per call rather than per token, so it needs its own line on your invoice and a bound per request, and it will be charged at what it costs us with nothing added.
What a Codex session costs
The same per-token prices as every other endpoint — see pricing. Two things are worth knowing before you point an agent at a small balance.
Credit is reserved up front. Codex sends no max_output_tokens, so we reserve against the model’s full output ceiling — around $0.25 per in-flight request on deepseek/deepseek-v4-flash. The unused part is refunded the moment the request finishes, so it is a floor on your balance rather than a cost, but a balance under a dollar will stop an agent mid-task with a 402.
Caching does most of the work. An agent resends a long instruction block every turn; those tokens come back as cache reads at a fraction of the input rate, itemised on each request in your dashboard. It is the single biggest difference between the bill you expect and the bill you get.
If something looks wrong
| Symptom | Cause |
|---|---|
| Reconnecting… 1/5, then a 402 | Out of credit. Codex retries a payment failure before reporting it, so one 402 looks like five. The message says how much the request needed. |
| Model metadata not found | Cosmetic. Codex asks /v1/models for its own catalogue format and gets OpenAI’s, which every other client expects. Nothing is degraded but the model picker. |
A 404 for a model you can see in the catalogue | The base_url is missing its /v1 suffix, so the request went somewhere else. Codex appends /responses to whatever it is given. |
| previous_response_id rejected | Conversations are not stored here. Send the full input each turn, which Codex does by default. |
Last updated 2026-09-28.