Skip to main content

config.toml in Codex CLI: The Complete Parameter Reference

Calculating read time…

config.toml is the single TOML file — ~/.codex/config.toml for your account, optionally layered with a .codex/config.toml per trusted project — that controls every dial OpenAI's Codex CLI exposes, from which model answers you to how hard it thinks, what it's allowed to touch, and which provider it talks to. Codex ships with an empty default file for a reason: sensible defaults cover most people, but the moment you want a cheaper model, a stricter sandbox, or a self-hosted endpoint, you're editing this file. 🗂️

Here's why getting the full picture matters, not just the five keys everyone copies from a README: some of the parameters people expect to find here — like a bare temperature or api_key field — don't work the way they assume, because Codex's native models are reasoning models tuned around effort levels rather than sampling temperature, and credentials are deliberately kept out of the TOML file entirely. Configure based on outdated assumptions and you'll either get an error, a setting that's silently ignored, or a key sitting in a file that shouldn't hold secrets. 🔐

Diagram showing config.toml organized into four groups: Model and Reasoning, Execution Safety, Provider and Auth, and Profiles

🔀 Quick Comparison — every parameter in this post
Parameter Purpose Typical value
modelDefault model for new sessions"gpt-5.5"
approval_policyWhen Codex must ask before acting"on-request"
sandbox_modeWhat Codex can technically touch"workspace-write"
model_reasoning_effortHow deeply the model thinks before answering"medium" (minimal/low/medium/high/xhigh)
tool_output_token_limitPer-tool-call output budget sent back to the model4000–12000
temperatureSampling randomness (non-reasoning providers only)Not used by native gpt-5.x reasoning models
max_output_tokens / model_max_output_tokensCaps response lengthModel-dependent override
env_key (not api_key)Names the env var Codex reads for auth"OPENAI_API_KEY"
base_urlCustom API endpoint, under [model_providers.id]Enterprise/self-hosted setups
--profileNamed configuration layer, selected at launchcodex --profile work

1. Where config.toml lives, and load order

User-level settings live at ~/.codex/config.toml (or $CODEX_HOME/config.toml if you've set that variable). A repository can add a .codex/config.toml, applied only while the project is trusted. When the same key appears in more than one place, priority goes: CLI flags and one-off -c key=value overrides, then project config, then an active profile file, then your user config, then any system-wide file, then built-in defaults. Codex ships an intentionally empty config.toml by default — every parameter below is opt-in.

2. model: picking the underlying intelligence

model = "gpt-5.5"
model_provider = "openai"

The model key sets the default for new sessions; model_provider points at which entry under model_providers (built-in openai, or a custom one) Codex should route requests through. Model naming has moved fast — Codex-tuned variants (historically suffixed -codex) sit alongside general-purpose models, and the fastest-moving part of any config guide is this exact field, so treat any specific model string, including the ones in this post, as something to confirm against the current model picker (/model in the TUI) rather than copy blindly from an old post.

🎯 Use this when: you want a stable, repo-wide default so you're not passing -m on every invocation.

3. approval_policy & sandbox_mode: the safety pair

approval_policy = "on-request"   # untrusted | on-request | never | granular
sandbox_mode = "workspace-write" # read-only | workspace-write | danger-full-access

approval_policy controls when Codex pauses to ask you; sandbox_mode controls what it's technically capable of doing, enforced at the OS level (Seatbelt on macOS, bwrap/seccomp on Linux). They're independent axes — "never ask" doesn't mean "no limits," it just means Codex won't prompt, while the sandbox still decides whether a write is even possible. on-request + workspace-write is the combination Codex recommends as its default "Auto" preset for version-controlled folders.

💡 Key distinction: approval_policy = "never" with sandbox_mode = "read-only" means Codex can inspect and explain code but cannot modify anything — not because it's asking politely, but because the sandbox blocks the write outright. We cover every practical combination of these two settings in a dedicated deep dive if you want the full matrix.

4. reasoning_effort, verbosity, and reasoning summaries

The full key is model_reasoning_effort, and current Codex builds support five levels, not three — minimal, low, medium (the default), high, and the newer xhigh for non-latency-sensitive work where quality matters more than speed:

model_reasoning_effort = "medium"      # minimal | low | medium | high | xhigh
plan_mode_reasoning_effort = "high"    # separate override, used only in /plan
model_reasoning_summary = "auto"       # auto | concise | detailed | none
model_verbosity = "medium"             # low | medium | high — output length/style

Higher effort means more internal reasoning tokens, slower responses, and higher cost, but materially better results on genuinely hard problems — reserve xhigh for the cases that actually need it rather than as a default. model_reasoning_summary controls how much of that thinking gets surfaced in the TUI; model_verbosity is a separate knob for how long and detailed the final response text itself is.

✅ Worked example: a team sets model_reasoning_effort = "medium" for day-to-day coding but plan_mode_reasoning_effort = "high", so quick edits stay fast while anything routed through /plan — where a bad plan compounds into a bad implementation — gets extra thought automatically.

5. tool_output_token_limit and context management

Every shell command, file read, or MCP response Codex runs produces output that gets stored in conversation history — an unbounded cat of a large log file can burn tens of thousands of tokens on its own. tool_output_token_limit caps that per call:

tool_output_token_limit = 4000
model_context_window = 200000          # override when Codex doesn't know a model's window
model_auto_compact_token_limit = 160000 # triggers automatic history summarization

When a tool's output would exceed the limit, Codex truncates or summarizes it before storing it, rather than letting one noisy command crowd out everything else in context. model_auto_compact_token_limit is the related control for the whole conversation: once history crosses that threshold, Codex runs a compaction pass to summarize older turns.

🎯 Use this when: you're running Codex against a repo with large log files, generated code, or verbose test output that otherwise dominates the context window.

6. temperature & max_output_tokens: the "if supported" parameters

This is the pair worth being precise about, because it's where copied config templates go stale fastest. Codex's own models are reasoning models, and reasoning-model APIs generally don't expose a sampling temperature the way older chat-completion models did — effort level and verbosity are the equivalent dials. temperature only becomes meaningful when you've pointed Codex at a custom, non-reasoning provider through model_providers that actually accepts it as a generation parameter:

[model_providers.custom]
name = "Custom Provider"
base_url = "https://api.example.com/v1"
env_key = "CUSTOM_API_KEY"
wire_api = "chat"

temperature = 0.2          # only meaningful for non-reasoning providers via wire_api
max_output_tokens = 8000   # response length cap; passed through to the provider
model_max_output_tokens = 100000  # Codex-native override for context accounting

In other words: for the built-in OpenAI reasoning models, skip temperature entirely and use model_reasoning_effort instead. If you're routing Codex through a custom or self-hosted model that isn't a reasoning model, temperature and max_output_tokens can be passed through as ordinary generation parameters — but confirm your provider actually reads them, since Codex forwards them without validating provider-side support.

💡 Contrasting case: setting temperature = 0.2 against the default OpenAI provider does nothing useful — the reasoning model ignores it. Setting it against a self-hosted Llama-family endpoint through a custom provider block genuinely changes output randomness.

7. api_key, env_key, and how authentication actually works

Codex deliberately does not have a plain api_key field in config.toml — putting a live secret in a plaintext config file that might get committed or shared is exactly the failure mode it avoids. Instead, each provider entry names an environment variable Codex should read at runtime:

[model_providers.openai]
env_key = "OPENAI_API_KEY"     # the variable NAME, not the secret itself

You export the real key in your shell (or a secrets manager), and Codex reads it from the environment at launch. For the default OpenAI provider, you can alternatively sign in interactively (codex login), which stores a ChatGPT-plan session instead of an API key — useful since usage on a ChatGPT plan is typically cheaper than metered API billing for regular, non-CI use.

✅ Worked example: a CI runner exports OPENAI_API_KEY as a masked secret in the pipeline, and its config.toml only ever contains env_key = "OPENAI_API_KEY" — the literal key value never touches version control or the config file at all.

8. base_url and custom model_providers

base_url doesn't sit at the top level — it belongs inside a named [model_providers.<id>] table, alongside the wire protocol Codex should speak to that endpoint:

[model_providers.internal_gateway]
name = "Internal Gateway"
base_url = "https://models.internal.company.com/v1"
env_key = "INTERNAL_MODEL_KEY"
wire_api = "responses"   # "responses" for OpenAI's newer API, "chat" for /chat/completions

model_provider = "internal_gateway"
model = "internal-model-id"

wire_api matters more than it looks: OpenAI's Responses API and the older Chat Completions protocol aren't interchangeable, and most third-party OpenAI-compatible gateways only implement the latter. Setting the wrong one is one of the most common reasons a "custom model isn't working" turns out to be a one-line fix.

🎯 Use this when: your organization runs models behind an internal gateway, Azure deployment, or a self-hosted inference server instead of talking to OpenAI directly.

9. profile: the (recently changed) way to switch setups

This is the parameter most likely to trip people up right now, because the mechanism changed. Older Codex versions supported a top-level profile = "name" selector paired with [profiles.name] tables inside the same config.toml. Current Codex builds have moved to standalone profile files instead — that in-file table syntax is deprecated and, in recent releases, no longer read at all:

# ~/.codex/config.toml — shared base
model = "gpt-5.5"
approval_policy = "on-request"
sandbox_mode = "workspace-write"

# ~/.codex/work.config.toml — only the values that differ
model_reasoning_effort = "high"
approval_policy = "untrusted"

# launch with:
# codex --profile work

A profile file only needs to hold the keys that differ from your base config — everything else falls through to ~/.codex/config.toml. If you're on an older Codex install still using [profiles.name] tables, plan to migrate: move each table's contents into its own <name>.config.toml file and drop the in-file table and the top-level profile selector.

💡 Why this changed: separate files make it obvious at a glance which profile you're loading (and let you diff them independently), instead of hunting through one large config.toml for the right nested table.

10. A complete, annotated config.toml

Putting every parameter from this post together into one realistic file:

# ~/.codex/config.toml

# ─── Model & reasoning ───
model = "gpt-5.5"
model_provider = "openai"
model_reasoning_effort = "medium"       # minimal | low | medium | high | xhigh
plan_mode_reasoning_effort = "high"
model_reasoning_summary = "auto"
model_verbosity = "medium"

# ─── Execution safety ───
approval_policy = "on-request"          # untrusted | on-request | never | granular
sandbox_mode = "workspace-write"        # read-only | workspace-write | danger-full-access

# ─── Context & tool output ───
tool_output_token_limit = 6000
model_context_window = 400000
model_auto_compact_token_limit = 200000

# ─── Custom provider (optional) ───
[model_providers.internal_gateway]
name = "Internal Gateway"
base_url = "https://models.internal.company.com/v1"
env_key = "INTERNAL_MODEL_KEY"          # NAME of an env var, never the key itself
wire_api = "responses"

Then, for a stricter reviewer setup, a separate ~/.codex/review.config.toml would only need the two or three lines that differ — say, approval_policy = "untrusted" and model_reasoning_effort = "high" — launched with codex --profile review.

11. Enterprise rollout at scale

Team-wide, three practices consistently work:

  1. Ship a locked base config, layer profiles on top. Keep organization-wide safety defaults (approval_policy, sandbox_mode, provider allowlist) in a managed base file, and let individuals add their own profile files for day-to-day speed rather than editing the shared one.
  2. Centralize provider credentials as env vars, never in TOML. Since Codex only ever reads an env_key name from the file, secret rotation becomes a secrets-manager operation, not a find-and-replace across a config repo.
  3. Tune tool_output_token_limit and context settings per workload, not per person. A CI profile reviewing large diffs needs a different budget than an interactive coding session; bake that into a shared ci.config.toml rather than leaving each engineer to discover the default is too small.
✅ Rollout pattern that scales: a platform team ships one base config.toml enforcing sandboxing and an approved provider list, publishes two ready-made profile files (fast-iterate and deep-review) engineers can opt into, and keeps every API key in the CI secrets store, referenced only by env_key name.

12. Common mistakes

  • Putting a real API key directly in config.toml. There's no api_key field to fill in on purpose — use env_key to name a variable and export the secret in your shell or secrets manager instead.
  • Setting temperature and expecting it to change anything. Against the default OpenAI reasoning models, it's a no-op; it only matters for custom, non-reasoning providers.
  • Putting base_url at the top level. It belongs inside a [model_providers.<id>] table, paired with env_key and the correct wire_api.
  • Copying an old [profiles.name] block from a pre-migration guide. Recent Codex versions read standalone profile files instead; an in-file profiles table may be silently ignored.
  • Leaving tool_output_token_limit at its default in a log-heavy repo. One verbose command can dominate context and degrade answer quality for the rest of the session — raise the limit deliberately for workloads that need it, rather than reactively after seeing bad output.

❓ FAQ

Q: Does config.toml have a literal api_key field?
A: No. It has env_key, which names an environment variable Codex reads at runtime — the actual secret never lives in the file.
Q: Why doesn't temperature seem to do anything?
A: Against OpenAI's built-in reasoning models it's not used — reasoning effort and verbosity are the equivalent controls. It only takes effect against custom, non-reasoning providers that accept it.
Q: Should I keep using [profiles.name] tables in config.toml?
A: For current Codex versions, no — move to standalone <name>.config.toml files selected with --profile; the in-file table syntax is deprecated and may not be read at all.
Q: What's the difference between tool_output_token_limit and model_context_window?
A: tool_output_token_limit caps a single tool call's output; model_context_window overrides the total token budget Codex assumes the active model has, which also drives when auto-compaction kicks in.
Q: How many reasoning effort levels does Codex actually support?
A: Five in current builds — minimal, low, medium, high, and xhigh — up from the three levels (low/medium/high) many older references still describe.
🔗 References & Further Reading

Codex and OpenAI are trademarks of OpenAI. This post explains and synthesizes publicly available documentation in its own words; it does not reproduce any source verbatim. Config keys, defaults, and profile mechanics change frequently — verify against the live docs and your installed Codex version before applying any config in production.

📝 Summary

  • config.toml loads with CLI flags > project config > profile file > user config > system config > defaults.
  • model / model_provider pick the intelligence; verify current model names before committing them to a shared config.
  • approval_policy and sandbox_mode are independent — one controls when Codex asks, the other what it can do.
  • model_reasoning_effort now spans five levels (minimal → xhigh), alongside model_verbosity and model_reasoning_summary.
  • tool_output_token_limit, model_context_window, and model_auto_compact_token_limit manage context pressure.
  • temperature and max_output_tokens are mostly relevant for custom, non-reasoning providers — not the default OpenAI models.
  • There's no api_key field by design — env_key names an environment variable instead.
  • base_url lives inside a [model_providers.<id>] table alongside wire_api.
  • Profiles are now standalone files selected with --profile, replacing the older in-file [profiles.name] tables.

That's the full map: a dozen or so keys, four control groups, and one file that quietly runs your whole Codex setup. Configure deliberately, and it stays predictable as your usage grows.

Comments