Skip to main content

Codex Hooks Explained: PreToolUse, PostToolUse, and SessionStart in OpenAI Codex CLI

Calculating read time…

Hooks in Codex are a built-in extensibility framework in OpenAI's Codex CLI that let you run your own deterministic scripts at fixed points in the agent's loop — before a tool runs, after it runs, when a session starts or ends, when the model's prompt is submitted, and more — so you can inspect, log, block, or rewrite what the agent is about to do without touching the model itself. 🪝

This matters because an autonomous coding agent that can run shell commands, edit files, and call MCP tools is only as safe as the guardrails wrapped around it. Without hooks, a team's only levers are prompt instructions (which a model can misread or ignore under pressure) and manual approval prompts (which don't scale past a handful of engineers). Hooks move policy enforcement — secret-key blocking, destructive-command interception, audit logging, compliance trails — out of the probabilistic model layer and into deterministic code that runs the same way every single time. That is the difference between "we told the agent not to do that" and "the agent structurally cannot do that." 🏛️

Diagram of the Codex agentic loop showing SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop, and SessionEnd hook injection points

The Codex agentic loop, annotated with every point where a hook can intervene.

🔀 Quick Comparison

Enforcement Method Where It Runs Can Block a Tool Call? Deterministic? Scales to a Team?
Codex Hooks Local script, outside the model Yes — deny, rewrite, or allow Yes Yes, via managed/requirements.toml
Prompt / AGENTS.md instructions Inside the model's context No — advisory only No Weak — model can drift
Manual approval prompts Human in the loop Yes, but per click Yes, but slow No — doesn't scale
Sandboxing alone OS / container boundary Only at the syscall boundary Yes Yes, but coarse-grained

1. What Are Hooks in Codex?

A Codex hook is a small, deterministic script that Codex invokes automatically at a defined moment in a conversation. Codex serializes the relevant context — session id, working directory, the tool about to be called, its arguments — into a single JSON object, writes it to the script's stdin, and reads the script's stdout, exit code, and stderr to decide what happens next. Nothing about that exchange touches the language model — the model never sees the hook's internal logic, and the hook doesn't need permission from the model to act.

Real, production-grade example — straight from OpenAI's own reference config

{
  "hooks": {
    "SessionStart": [
      { "matcher": "startup|resume",
        "hooks": [{ "type": "command",
          "command": "python3 ~/.codex/hooks/session_start.py",
          "statusMessage": "Loading session notes" }] }
    ],
    "PreToolUse": [
      { "matcher": "Bash",
        "hooks": [{ "type": "command",
          "command": "python3 .codex/hooks/pre_tool_use_policy.py",
          "statusMessage": "Checking Bash command" }] }
    ]
  }
}

Read this line-by-line the way a platform engineer would:

  1. "hooks" is the top-level container — everything under it is organized by event name, so Codex can look up "what fires right now" in constant time.
  2. "SessionStart" is an event key. Its matcher of "startup|resume" is a regular expression tested against how the session began — so this handler skips silently on a compact-triggered restart, which needs different logic.
  3. The nested "hooks" array under each matcher lets you attach more than one handler to the same trigger — Codex runs all of them.
  4. "PreToolUse" with matcher "Bash" is the enterprise workhorse: every shell command the agent wants to run passes through this script first.
  5. "statusMessage" is what a developer sees in the CLI while the hook runs — without it, a slow policy check just looks like a frozen terminal, and support tickets follow.

What breaks in production without this: an agent with shell access and no PreToolUse policy hook will, sooner or later, run something a human reviewer would have stopped — not out of malice, but because a multi-step plan drifted and a destructive command looked locally reasonable to the model in that context window. Hooks are the seatbelt, not the airbag; they stop the bad command before it executes, not after.

✅ Worked example: Tie this back to the config above — if pre_tool_use_policy.py sees tool_input.command containing rm -rf against a path outside the sandboxed workspace, it exits with code 2 and writes the reason to stderr. Codex denies the call before the shell ever spawns.

🎯 Use this when you need policy enforcement that survives a bad prompt, a confused plan, or an adversarial input — not just when the model is behaving well.

2. The Agentic Loop: Where and When Hooks Fire

Codex documents its hook surface as covering three broad moments: during a turn (PreToolUse, PermissionRequest, PostToolUse, PreCompact, PostCompact, UserPromptSubmit, SubagentStop, Stop), when a session or subagent starts (SessionStart, SubagentStart), and when the main thread ends (SessionEnd, which does not fire for subagents). That three-way split matters operationally: turn-level hooks are where you enforce policy on individual actions, while session-level hooks are where you handle onboarding context and cleanup.

A concrete, verifiable production pattern that OpenAI's own documentation calls out directly: teams use UserPromptSubmit to scan a prompt for an accidentally pasted API key before it ever reaches the model, and PreToolUse to run a custom validation check that enforces coding standards before a shell command executes. Neither of those checks depends on the model choosing to comply — they run whether the model "wants to" or not.

💡 Contrast this with: a purely prompt-based instruction like "never commit secrets." That instruction lives in the model's context window and competes for attention with everything else in a long session. A UserPromptSubmit hook that regex-scans for key patterns runs every single time, with zero dependency on context-window pressure.

🎯 Use this when you're deciding which lifecycle stage owns a given policy — turn-level hooks for per-action gates, session-level hooks for onboarding and teardown.

3. Hook Anatomy: Events, Matchers, and Handlers

Every hook definition has three layers, and understanding them is what separates a hook that fires correctly from one that silently never runs:

  1. Event — the lifecycle moment, such as PreToolUse or SessionStart.
  2. Matcher — a regex string filtering when that event actually triggers the handler. Leaving matcher empty, or setting it to "*", matches every occurrence of a supported event. Not every event honors a matcher — UserPromptSubmit and Stop ignore any matcher you configure.
  3. Handler — the actual command to run, with an optional timeout (600 seconds by default for most hooks; 1 second, extendable to 3, for SessionEnd) and an optional statusMessage.

Codex discovers hooks from several config layers and merges all matching ones rather than letting a higher-precedence layer silently override a lower one — a personal ~/.codex/hooks.json, a repo-local .codex/hooks.json, inline [hooks] tables in config.toml, and hooks bundled with an installed plugin's manifest all load together. If two matching command hooks fire on the same event, they run concurrently — one can't block another from starting, which is an important constraint if you're building hooks that assume ordering.

Real example — the same policy hook, written in inline TOML

[[hooks.PreToolUse]]
matcher = "^Bash$"

[[hooks.PreToolUse.hooks]]
type = "command"
command = '/usr/bin/python3 ".codex/hooks/pre_tool_use_policy.py"'
timeout = 30
statusMessage = "Checking Bash command"

Line-by-line: the double-bracket [[hooks.PreToolUse]] syntax is TOML's way of appending an array entry — each block is one matcher group. The regex anchors (^Bash$) matter for teams running multiple matchers on similar tool names; an unanchored Bash would also match a hypothetical future tool named BashLegacy. The explicit timeout = 30 caps how long a slow policy check can hang the turn — without it, the 600-second default silently swallows a hung script for ten minutes before anyone notices.

🎯 Use this when you're auditing why a hook you configured "isn't firing" — the matcher regex is the first thing to check, then which config layer actually loaded it.

4. PreToolUse: Blocking and Rewriting Tool Calls

PreToolUse is the highest-leverage hook in the whole framework because it runs before the tool executes and can change the outcome three different ways: deny the call outright, allow it with rewritten arguments, or allow it unchanged while injecting extra context for the model. It covers shell commands, unified-exec sessions, file edits made through apply_patch, MCP tool calls, and most other local function tools — though hosted tools like web search don't route through this path, since they never touch the local function-tool layer.

Diagram showing PreToolUse hook's three possible responses: deny, allow and rewrite, or annotate only

A PreToolUse hook has exactly three exits: deny, allow-with-rewrite, or annotate.

Real example — the deny response, straight from the wire format

{
  "hookSpecificOutput": {
    "hookEventName": "PreToolUse",
    "permissionDecision": "deny",
    "permissionDecisionReason": "Destructive command blocked by hook."
  }
}
  1. The script receives tool_name (for example Bash or an MCP name like mcp__fs__read) and tool_input on stdin.
  2. It inspects tool_input.command for shell calls, or the tool-specific argument object for MCP calls.
  3. To deny, it prints the JSON above to stdout — or, more simply, exits with code 2 and writes the reason to stderr.
  4. To rewrite instead of deny, it returns permissionDecision: "allow" together with an updatedInput object; for Bash and apply_patch that means a replacement command string.

What breaks without this: teams that only use PostToolUse for governance discover, the hard way, that reviewing after execution can't undo a shell command that already ran. PreToolUse is the only hook in the framework that can prevent a side effect from happening at all.

💡 Key warning: continue, stopReason, and suppressOutput are not supported fields for PreToolUse today. If a hook returns one of them, Codex marks that hook run as failed, reports the error, and lets the tool call continue anyway — the opposite of the fail-safe behavior most teams assume. Test your hook's exact output shape, not just its logic.

🎯 Use this when the cost of a tool call executing even once is unacceptable — credential exposure, irreversible deletes, production database writes.

5. PostToolUse and PermissionRequest

PostToolUse runs after a supported tool produces output — including after a Bash command exits with a non-zero status — and it can review, log, or replace what the model sees, but it cannot undo whatever side effect already happened. Its decision: "block" response doesn't roll back the command; instead, Codex swaps the tool result for the hook's feedback message and continues the model from there. That's a subtle but critical distinction from PreToolUse's deny behavior.

PermissionRequest sits at a different point: it fires when Codex is about to surface an approval prompt to a human — a shell escalation, for instance — and it can resolve that prompt automatically by returning an allow or deny decision, or decline to decide and let the normal prompt appear. If multiple matching hooks return conflicting decisions, any deny wins over an allow — a deliberately conservative default.

✅ Worked example: continuing the Bash policy script from Section 1 — a matching PostToolUse handler can pipe the same command's stdout through a secrets-scanner. If it finds a leaked token in the output, it returns decision: "block" with a reason, and the model receives that reason instead of the raw output containing the secret.

🎯 Use this when you need an audit trail or a last line of defense on output content — not when you need to stop a call from running at all.

6. Session Lifecycle: SessionStart, SessionEnd, Compaction

SessionStart fires when a session begins, resumes, is cleared, or restarts after context compaction — and its matcher filters on which of those four sources triggered it. Plain text a script prints to stdout is added directly as developer context; JSON output can set hookSpecificOutput.additionalContext for the same purpose with more structure. This is the standard mechanism teams use to inject house style guides, recent incident notes, or repo-specific conventions automatically, without asking every engineer to paste them manually.

SessionEnd is advisory only — it runs on session close, archive, delete, or a 30-minute idle timeout, but never for subagents, and its output can't keep the thread open or steer the model. It exists for cleanup: writing final notes, releasing a lock file, or flushing a metrics buffer. PreCompact and PostCompact bracket Codex's own history-compaction step, each filterable by whether compaction was manual or auto.

💡 Contrast this with PreToolUse: where PreToolUse governs a single action, SessionStart governs the entire session's starting context. A common mistake is stuffing large amounts of text into SessionStart's additionalContext — Codex caps model-visible hook output at roughly 2,500 tokens per handler by default (configurable via additionalContextLimit), spilling anything larger to disk and showing the model only a preview.

🎯 Use this when onboarding context needs to reach the model automatically at the start of every session, rather than living in a document nobody re-reads.

7. UserPromptSubmit, Stop, and Subagent Hooks

UserPromptSubmit fires on every prompt right before Codex sends it to the model, and it's the natural place for input-side data-loss prevention: scanning for API keys, PII, or other sensitive strings before they enter the model's context at all. Stop fires when a turn is about to end and can return decision: "block" to keep the agent going — Codex treats the hook's reason text as a new continuation prompt, effectively saying "not done yet, try this next." That makes Stop the mechanism behind automated "keep working until tests pass" workflows.

SubagentStart and SubagentStop mirror SessionStart and Stop but for delegated subagents, each carrying an agent_type field the matcher can filter on. This lets a platform team apply a stricter policy to a subagent spawned for, say, dependency upgrades than to the primary agent handling day-to-day edits.

✅ Worked example: a Stop hook that runs the test suite. If tests fail, it returns {"decision": "block", "reason": "Two tests still fail in checkout.spec.ts — fix them before finishing."}, and Codex automatically feeds that back as the next prompt — no human has to notice and re-prompt manually.

🎯 Use this when "done" needs a machine-checkable definition — test suites passing, lint clean, a specific file present — rather than trusting the model's own judgment that it's finished.

8. Trust, Security, and Managed Hooks

Because a hook is arbitrary code that runs automatically, Codex requires explicit review before most hooks can execute. Non-managed command hooks must be reviewed and trusted — Codex records that trust against a hash of the hook's exact current definition, so any edit to the script or its config marks it for re-review and skips it until trusted again. The /hooks slash command in the CLI is where a developer inspects sources, reviews new or changed hooks, and disables individual ones.

Enterprises get a separate, higher tier: hooks defined under [hooks] in a requirements.toml file, or delivered via MDM/system/cloud channels, are marked managed — trusted by policy and not removable from the individual user's hook browser. Admins can additionally set allow_managed_hooks_only = true to ignore every user, project, session, and plugin hook, running only the centrally managed set — and can pin [features].hooks = true in the same file so a user can't disable hooks locally even if they wanted to.

Real example — enterprise-managed hook enforcement

allow_managed_hooks_only = true

[features]
hooks = true

[hooks]
managed_dir = "/enterprise/hooks"

[[hooks.PreToolUse]]
matcher = "^Bash$"

[[hooks.PreToolUse.hooks]]
type = "command"
command = "python3 /enterprise/hooks/pre_tool_use_policy.py"
timeout = 30

Line-by-line, for the platform team deploying this: allow_managed_hooks_only is the kill switch that stops a rogue project-local .codex/hooks.json from ever taking priority; managed_dir tells Codex where to expect scripts on disk, but Codex itself does not distribute or update those files — an MDM or configuration-management pipeline still owns delivery. Skipping that last part is the single most common gap in first-pass enterprise rollouts: admins set the policy in requirements.toml and forget that the actual script still has to land on every machine through a separate channel.

💡 Key warning: for one-off automation that has already vetted its hook sources outside Codex — a CI pipeline, for instance — the --dangerously-bypass-hook-trust flag skips the persisted trust check for that single invocation. The name is not decorative; using it in an interactive developer session defeats the entire review model this section describes.

🎯 Use this when a policy must hold regardless of individual developer configuration — regulatory requirements, SOC 2 controls, or any control an auditor will ask about by name.

9. Enterprise Rollout at Scale

Rolling hooks out across an engineering org is a governance problem as much as a technical one. A pattern that holds up in practice:

  1. Own the taxonomy first. Decide, before writing scripts, which policies are non-negotiable (managed, via requirements.toml) versus which are team-level defaults developers can override (project .codex/hooks.json, checked into the repo like any other config).
  2. Template the handler script, not the policy. Give every team a shared library for reading Codex's stdin JSON and emitting the correct hookSpecificOutput shape, so individual policy scripts stay short and reviewable rather than each reinventing the wire protocol.
  3. Enforce in CI, not just at the desktop. Codex's non-interactive mode and CI integrations honor the same hook config — treat a missing or disabled managed hook in a pipeline run as a failed build, the same way you'd treat a missing lint config.
  4. Version and code-review hook scripts like production code. Because Codex re-triggers the trust review on any hash change, a pull-request workflow for hook scripts gives you both a security review and an automatic Codex re-trust prompt for free.
  5. Instrument metrics from day one. A PostToolUse or Stop hook is a natural place to emit structured logs — blocked-command counts, average hook latency, false-positive rate on secret scanning — to whatever observability stack the org already runs, so the security team can tune thresholds instead of guessing.

✅ Worked example: the managed pre_tool_use_policy.py from Section 8 gets a companion PostToolUse handler that increments a Prometheus counter every time it fires a deny — giving the security team a live dashboard of exactly which commands get blocked most often, without reading raw logs.

🎯 Use this checklist when hooks move from "one engineer's personal safety net" to "the org's control plane for agentic coding."

🚫 Common Mistakes

1. Treating PostToolUse as if it can undo a side effect. It can't — the tool already ran. If the goal is prevention, the logic belongs in PreToolUse, not after the fact. This is the single most common architecture mistake, because reviewing output feels like the natural place to add "safety" logic.

2. Assuming Bash-only coverage. Older mental models of Codex hooks (and some blog posts) describe PreToolUse as shell-only. Current Codex hook coverage extends to apply_patch, MCP tools, and other local function tools — teams that only wrote a Bash matcher leave file edits and MCP calls completely ungoverned.

3. Returning unsupported output fields and assuming fail-closed behavior. If a PreToolUse hook returns continue or stopReason — fields that event doesn't support — Codex logs an error and lets the tool call proceed anyway. Silence is not safety; test the exact output contract for each event.

4. Forgetting that concurrent hooks can't order themselves. Multiple matching command hooks for one event launch concurrently, so a script that assumes "my logging hook runs before the policy hook" has no guarantee of that — design each hook to be independent, not sequential.

5. Shipping managed policy without shipping the script. Setting managed_dir in requirements.toml only tells Codex where to look — it doesn't install anything. Teams that skip the separate MDM/config-management delivery step end up with a policy that silently never runs.

6. Returning secrets in hook output. Oversized hook output gets spilled to disk under a temp directory for the model's preview mechanism — a hook that echoes back the very secret it just scanned for defeats its own purpose and leaves that secret sitting in a file.

❓ FAQ

Are hooks enabled by default in Codex?

Yes. Hooks ship enabled by default; a team can turn them off locally by setting [features].hooks = false in config.toml, though admins can override that with a managed requirements.toml setting.

How is a Codex hook different from an AGENTS.md instruction?

An AGENTS.md instruction lives inside the model's context and is advisory — the model can misread or deprioritize it. A hook is a separate script that runs deterministically outside the model, with the power to actually deny a tool call rather than merely ask the model not to make it.

Can a PreToolUse hook block an MCP tool call?

Yes. MCP tool calls route through the same local function-tool hook path as shell commands and file edits, so a matcher like mcp__filesystem__.* can deny or rewrite them the same way it would a Bash command.

Do hooks run for subagents the same way they run for the main session?

Mostly, but not entirely. Turn-level hooks like PreToolUse apply to subagent tool calls too, and subagents get their own SubagentStart/SubagentStop pair — but SessionEnd specifically does not run for subagents, only for the main thread.

Can a hook see and leak the model's full conversation?

A hook receives a transcript_path pointing to the session's transcript file, so a script could read it — which is exactly why enterprise rollouts treat hook scripts as production code subject to review, not casual personal automation, and why hook output itself should never echo back sensitive data.

🔗 References & Further Reading

All product names, trademarks, and registered trademarks (including Codex and OpenAI) are the property of their respective owners and are referenced here for identification and educational purposes only. Every explanation above is original synthesis written to fact-check against the sources listed — no text is reproduced verbatim from any source.

📝 Summary

  • Hooks are deterministic scripts Codex runs at fixed lifecycle moments, outside the model's own reasoning.
  • The agentic loop exposes turn-level hooks (PreToolUse, PostToolUse, PermissionRequest, Stop, and more) and session-level hooks (SessionStart, SessionEnd).
  • Every hook config has three layers: event, matcher regex, and command handler.
  • PreToolUse is the only hook that can prevent a tool call from running at all — deny, rewrite, or annotate.
  • PostToolUse and PermissionRequest govern review and approval, not prevention.
  • Session lifecycle hooks handle onboarding context and cleanup, bracketing compaction too.
  • UserPromptSubmit, Stop, and subagent hooks extend governance to prompts and delegated work.
  • Enterprises get a managed tier via requirements.toml, with hash-based trust review for everything else.
  • Rollout at scale is a governance exercise: taxonomy, templated scripts, CI enforcement, and metrics.
  • The most common mistakes all trace back to one misunderstanding: which hook can actually prevent an action, and which can only react to one.

That's the full picture of hooks in Codex — from the JSON hitting a script's stdin to the managed policy an enterprise admin pins across an entire org. Go configure one small PreToolUse guard today; it's the fastest way to feel the difference between an agent you trust because it behaved well so far, and one you trust because it structurally can't do the thing you're worried about. 🛠️

Comments