Skip to main content

The Codex CLI Playbook: 9 Repeatable Workflows Every Enterprise Team Should Standardize On

Calculating read time…

Codex's documentation formalizes nine repeatable workflow recipes — Explain, Fix, Test, Prototype, UI Iterate, Delegate, Local Review, PR Review, and Update Docs — and every one of them is really the same four-part contract wearing different clothes: a Goal, some Context, a set of Constraints, and a Done-when condition that tells Codex (and you) when to stop. Learn the skeleton once, and you can improvise a tenth recipe for whatever your team actually does. 🧭

Here's why this is worth internalizing rather than just bookmarking: each recipe also recommends a specific Codex surface — the IDE extension, the CLI, a delegated cloud task, or a GitHub comment — because the right amount of context-gathering and the right amount of autonomy aren't the same for "explain this function" as they are for "refactor the auth subsystem." Use the wrong surface for the wrong risk level, and you either babysit something that should have run unattended, or you unleash something that needed a tighter leash. ⚙️

Diagram showing the nine Codex workflow recipes grouped by surface — IDE/CLI, Cloud, and GitHub — all sharing the Goal, Context, Constraints, Done-when skeleton

🔀 Quick Comparison — all nine recipes at a glance
# Recipe Best surface Risk / scope
1Explain a codebaseIDE or CLIRead-only, single area
2Fix a bugCLI (or IDE)Write, narrow scope
3Write a testIDE (selection-based) or CLIWrite, single function
4Prototype from a screenshotCLI or IDE (image input)Write, new files
5Iterate on UI liveCLI + running dev serverWrite, tight loop
6Delegate a refactorIDE plan → Cloud executeWrite, multi-file
7Local code reviewCLI (/review)Read-only, working tree
8PR reviewGitHub commentRead-only, async
9Update docsIDE or CLIWrite, low risk

The shared skeleton behind all nine recipes

Every recipe on Codex's workflow page reduces to the same four questions: what outcome are you after (Goal), what should Codex look at to get there (Context), what's off-limits while doing it (Constraints), and how will you both know it's finished (Done-when — the same idea we called "Completion criteria" in our prompt-anatomy deep dive). The recipes differ in which of the four gets emphasized: a bug fix leans hard on Context (the reproduction steps matter more than the description), while a documentation update leans on Done-when (a broken link is a very checkable failure). Treat each recipe below as that same skeleton, tuned for its situation.

1. Explain a codebase

Use it when: onboarding onto a service, inheriting unfamiliar code, or needing to understand a data flow before you touch it. Best on the IDE extension for fast local exploration, or the CLI when you want a saved transcript.

Goal: Understand how webhook payloads are verified before being processed.
Context: @webhook_handler.py @signature.py — focus on where the signature
is checked and where a bad signature is rejected.
Constraints: Explain only; do not propose changes yet.
Done-when: I can describe, in order, every check a payload passes through
before it reaches business logic.
✅ Worked example: asking Codex to also produce a numbered checklist of the steps, not just prose, turns the explanation into something you can verify against the actual code in thirty seconds — a much faster check than re-reading a paragraph.

🎯 Use this when: the Goal is understanding, not editing — resist the urge to also ask for a fix in the same prompt; that's recipe 2.

2. Fix a bug

Use it when: you have a reproducible failure. Best on the CLI, since Codex benefits from running your reproduction steps itself rather than trusting a description of the bug.

Goal: Fix duplicate notification emails sent by the digest worker.
Context: Repro — run the worker locally, enqueue two events for the same
user within one second, and two identical emails are sent instead of one.
Constraints: Do not change the public event schema. Keep the fix minimal;
add a regression test if reasonably possible.
Done-when: Running the repro sends exactly one email, and the existing
worker test suite still passes.
💡 What matters most here: the reproduction steps, not the bug description. "Emails sometimes duplicate" gives Codex nothing to run; a numbered repro gives it something to reproduce, fix, and re-verify against — the same loop you'd want from a human engineer.

3. Write a test

Use it when: you want a test scoped to a specific function rather than a vague "add more tests" request. Selection-based in the IDE (select the function, add it to the thread), or path-based on the CLI.

Goal: Add unit tests for the mergeIntervals utility.
Context: @intervals.ts — follow the naming and assertion style already
used in the neighboring test file for sortByStart.
Constraints: No new testing libraries.
Done-when: Tests cover the happy path, an already-merged input, and an
empty array, and the full test file passes.

🎯 Use this when: the function's behavior at its edges is exactly what you're unsure about — naming those edges in Constraints or Done-when is more useful than a generic "add tests" ask.

4. Prototype from a screenshot

Use it when: you have a design mock or reference image and want a working first pass fast. The CLI accepts a dragged-in image file directly in the prompt; the IDE extension accepts drag-and-drop or paste in chat.

Goal: Build a working prototype of the attached account-settings mockup.
Context: [attached image] — match layout, spacing, and type scale closely.
Constraints: React, Vite, Tailwind, TypeScript. Reuse existing form
components where they already exist in the project.
Done-when: The page renders at /settings, the dev server runs, and the
layout visually matches the mockup at desktop width.
💡 What the image doesn't tell Codex: hover states, validation rules, and keyboard behavior aren't visible in a static screenshot. Spell those out in text, or expect a prototype that looks right but behaves like a static mock.

5. Iterate on UI with live updates

Use it when: you want a design → tweak → refresh loop rather than a one-shot prototype. Run your dev server in a separate terminal so you can watch changes land in the browser as Codex edits.

Goal: Improve the visual hierarchy of the pricing table header only.
Context: Dev server already running at localhost:5173/pricing.
Constraints: Touch only the header section; leave the plan cards
untouched for this pass.
Done-when: I confirm the header looks right in the browser, then say
"looks good" before you move to the next section.

🎯 Use this when: you plan to accept or reject each small change visually — keep each request scoped to one section so a rejected change doesn't undo something you already liked elsewhere.

6. Delegate a refactor to the cloud

Use it when: the work is well-understood but long — plan it carefully with fast local context, then hand the actual implementation to an isolated cloud task so it can run without blocking you.

Goal: Split the monolithic OrderService into order-creation,
order-pricing, and order-fulfillment modules.
Context: Plan locally first, scanning current module boundaries and
call sites. Then delegate implementation of each milestone to a cloud task.
Constraints: No changes to the public OrderService API during the
migration; each milestone must leave the build green.
Done-when: Each milestone's cloud diff compiles, existing tests pass,
and a PR is opened for review.
✅ Worked example: asking Codex to produce a milestone-by-milestone plan locally first — with a rollback note per milestone — before delegating gives you a checkpoint to negotiate scope before any cloud task starts making changes you'd have to unwind.

7. Do a local code review

Use it when: you want a second set of eyes on your working tree before committing or opening a PR. This is largely a single command rather than a prompt you have to construct yourself.

/review
/review Focus on input validation and auth-boundary issues

🎯 Use this when: you're about to open a PR — apply the fixes it surfaces, then re-run /review to confirm the issues are actually resolved rather than assuming they are.

8. Review a GitHub pull request

Use it when: you want review feedback without pulling the branch locally, or you want a reviewer that's always available for teammates' PRs. Requires Codex code review enabled on the repository.

@codex review
@codex review for authentication and authorization edge cases

The Done-when here is implicit but real: the review comment itself is the deliverable, and "done" means the comment addresses the focus area you named, not just a generic pass.

9. Update documentation

Use it when: a doc needs to change and you want the result checked, not just drafted. Works equally well from the IDE or CLI.

Goal: Update the rate-limiting doc page to describe the new
per-API-key limits and the 429 response format.
Context: @docs/rate-limiting.md — follow the existing heading structure.
Constraints: Do not remove the legacy IP-based limiting section; mark it
deprecated instead.
Done-when: All links in the page resolve, and the new limits match what
the /api/search endpoint actually returns.
💡 Easy to skip, worth keeping: asking Codex to verify links and cross-check the documented behavior against the actual code catches the most common doc-update failure — a docs change that's well-written but quietly wrong.

Chaining recipes across one real session

A realistic feature session rarely uses one recipe in isolation. A common chain: Explain the affected module first, Fix or extend it, Write a test for the change, run Local code review before committing, open a PR, and let PR review catch anything a human reviewer might otherwise flag first. For anything that grows beyond an afternoon, that's the moment to switch from doing it inline to Delegate-ing the remaining milestones to a cloud task while you move on.

Enterprise rollout at scale

Three practices make these recipes stick across a team rather than living in one engineer's muscle memory:

  1. Turn the recipes into snippets, not tribal knowledge. Keep a short library of Goal/Context/Constraints/Done-when templates per recipe — bug fix, refactor plan, doc update — so new team members reach for a proven structure instead of writing prose from scratch.
  2. Standardize on /review before PR, @codex review after. Running the same reviewer locally and on GitHub means issues get caught earlier without adding a second, different review standard.
  3. Reserve cloud delegation for planned, multi-milestone work. Ad hoc delegation of poorly scoped tasks produces diffs nobody wants to review; a local plan first (recipe 6) keeps cloud tasks accountable to something concrete.

Common mistakes

  • Using "Fix a bug" without a reproduction. A description of symptoms isn't a repro; Codex (and any engineer) needs steps that reliably trigger the failure to verify a fix actually worked.
  • Skipping the local plan before delegating to the cloud. Cloud tasks run in isolation and can drift further from what you wanted with no one watching — plan the milestones first.
  • Treating a screenshot as a complete spec. A static image can't show hover states, validation, or keyboard behavior — say those out loud in Constraints or they won't exist in the prototype.
  • Bundling UI iteration requests across multiple sections at once. Small, section-scoped prompts let you accept or reject each change independently; a single sweeping request makes partial rollbacks messy.
  • Treating PR review as a substitute for local review. They catch overlapping but not identical issues; run /review before opening the PR rather than relying on @codex review alone.

❓ FAQ

Q: Do I need to write out Goal/Context/Constraints/Done-when explicitly every time?
A: Not word-for-word, but the four ideas should be present. "/review" is a one-word Goal because Codex already knows the Context (your working tree) and a sensible Done-when (issues resolved) by default — explicit structure matters most when any of the four would otherwise be ambiguous.
Q: How do I choose between the CLI and the IDE extension for a given recipe?
A: The IDE extension automatically includes your open files as context, which suits fast, visual, selection-based tasks like writing a test or iterating on UI. The CLI is better when you want an explicit transcript, want to attach specific files with @, or are running longer, less visual work.
Q: When should a task go to the cloud instead of running locally?
A: When it's well-scoped but long enough that you don't want to babysit it — a multi-milestone refactor is the canonical case. Plan it locally first so the cloud task has a concrete target rather than an open-ended goal.
Q: What's the difference between local review and PR review?
A: Local review (/review) checks your working tree before you commit or open a PR. PR review (@codex review) runs against an already-opened GitHub pull request, useful for reviewing without pulling the branch, or for reviewing teammates' PRs.
Q: Can I invent a tenth recipe for something specific to my team?
A: Yes — that's the point of learning the underlying skeleton rather than memorizing nine fixed prompts. Any repeatable task your team does regularly (a deploy checklist, a data-migration pattern) can follow the same Goal/Context/Constraints/Done-when shape.
🔗 References & Further Reading

Codex and OpenAI are trademarks of OpenAI. This post explains and synthesizes publicly available documentation in its own words, with original examples rather than reproduced ones; it does not reproduce OpenAI's documentation verbatim. Workflow details and available commands may change — verify against the live docs before relying on them in production.

📝 Summary

  • All nine recipes share one skeleton: Goal, Context, Constraints, Done-when.
  • Explain / Fix / Test / Prototype / UI Iterate are fast, local loops — best on IDE or CLI.
  • Delegate is for well-planned, longer work — plan locally, execute in an isolated cloud task.
  • Local review checks your working tree; PR review checks an already-open pull request.
  • Update docs is low-risk but benefits from an explicit, checkable Done-when just like any other recipe.
  • A realistic session chains several recipes together rather than using one in isolation.
  • The nine recipes are examples of the pattern, not the limit of it — build your own from the same skeleton.

Nine names, one underlying contract. Learn the shape once, and every new situation just becomes a matter of filling in four blanks.

Comments