The Anatomy of an Effective Codex Prompt: Goal, Context, Constraints, Completion
An effective Codex CLI prompt is really a tiny spec, and reliably good ones share four ingredients: a Goal, Context, Constraints, and Completion criteria. Miss one, and Codex either guesses at what you meant, picks its own libraries and patterns, or stops halfway through because it isn't sure what "done" means. 🎯
Here's why this matters more than a style preference: Codex works in a loop, calling the model and then acting on file reads, edits, and commands until it decides the task is finished. If your prompt doesn't define completion, Codex has to invent its own finish line — and an invented finish line is exactly where "close enough" implementations, half-applied refactors, and silently skipped edge cases come from. A well-structured prompt spends your Codex usage on implementation instead of clarification. 🧩
- Why Codex needs this structure at all
- Goal: what "done" looks like
- Context: what Codex should look at first
- Constraints: what Codex must never do
- Completion criteria: the check that ends the loop
- The full worked example, line by line
- From one-off prompts to Goal mode
- Enterprise rollout: AGENTS.md and shared templates
- Common mistakes
- FAQ
| Element | Answers | Missing it causes |
|---|---|---|
| Goal | What outcome am I asking for? | Vague, open-ended implementations |
| Context | What should Codex look at first? | Off-pattern code that ignores your codebase's conventions |
| Constraints | What must Codex avoid? | New dependencies, touched schemas, or edited files you didn't want changed |
| Completion criteria | How does Codex know it's finished? | Work that stops halfway, or continues past "good enough" |
1. Why Codex needs this structure at all
Codex isn't a single-shot autocomplete — it's an agent running a loop of model calls and tool actions (file reads, edits, shell commands) until it judges the task complete. That loop needs an exit condition, and if you don't supply one, the model has to infer it. A prompt like "make this better" or "refactor this code" gives Codex no reliable finish line, so it either stops too early on a convenient proxy (like "the tests I happened to touch still pass") or keeps making changes you never asked for. The four-element structure exists to remove that ambiguity before the loop even starts.
2. Goal: what "done" looks like
The Goal is one concrete sentence describing the observable outcome, not the steps to get there. Use outcome verbs — add, fix, migrate, replace — rather than process verbs like improve or refactor, which describe an activity but not an end state.
Goal: Add per-API-key rate limiting to the /api/search endpoint.
3. Context: what Codex should look at first
Context points Codex at the language version, framework, and — critically — an existing pattern in your own codebase to mirror, rather than whatever pattern the model would default to on its own. Without this, Codex may quietly choose a different library, a different throttling strategy, or a different response format than the rest of your project uses, simply because nothing told it not to.
Context: Python FastAPI service, Redis already used for caching, follows the token-bucket pattern already implemented in src/api/auth.py for throttling failed login attempts.
Every time you find yourself copy-pasting the same background paragraph into prompt after prompt, that's a signal it belongs in a persistent AGENTS.md file at the repo root instead — more on that in the enterprise section below.
4. Constraints: what Codex must never do
Constraints are the boundaries: forbidden dependencies, protected files, schema changes that are off the table. This is the element people skip most often, and it's the one that causes the most expensive mistakes, because an agent that can edit files and run commands will happily "solve" a problem in a way you didn't want if nothing told it not to.
Constraints: Do not add new dependencies (redis client is already in requirements.txt). Do not change the response schema for successful requests. Do not modify the /api/auth endpoint.
requirements.txt, instead of quietly pulling in a separate rate-limiting library nobody reviewed.
5. Completion criteria: the check that ends the loop
This is the exit condition for the agent's loop, and it should be something concrete and — ideally — runnable, not a subjective judgment call. "The code looks good" isn't verifiable. "Requests beyond the configured limit return HTTP 429, and pytest exits with code 0" is something both you and Codex can check.
Completion: A client sending more than 10 requests per minute to /api/search receives a 429 response with a Retry-After header on the 11th request. Existing tests still pass. Run pytest.
🎯 Use this when: the task has any ambiguity at all about what "finished" means — which, in practice, is almost every non-trivial task.
6. The full worked example, line by line
Put together, the four elements read as a single, unambiguous prompt:
Goal: Add per-API-key rate limiting to the /api/search endpoint. Context: Python FastAPI service, Redis already used for caching, follows the token-bucket pattern already implemented in src/api/auth.py for throttling failed login attempts. Constraints: Do not add new dependencies (redis client is already in requirements.txt). Do not change the response schema for successful requests. Do not modify the /api/auth endpoint. Completion: A client sending more than 10 requests per minute to /api/search receives a 429 response with a Retry-After header on the 11th request. Existing tests still pass. Run pytest.
Notice what each line buys you: the Goal stops Codex from wandering into unrelated refactors; the Context stops it from inventing a new throttling strategy instead of reusing the token-bucket approach already proven in auth.py; the Constraints stop it from adding a dependency or touching an endpoint you didn't ask about; and the Completion criteria give it — and you — an objective way to know the run is over. One specific example like this is worth more in a prompt than a paragraph of general explanation.
7. From one-off prompts to Goal mode
The four-element structure works well for single-session tasks, but longer, multi-step work benefits from Codex's built-in Goal mode, started with /goal (enable it first with features.goals = true in config.toml if it isn't already on). A goal acts as both the starting prompt and the persistent completion criteria Codex checks against as it works across many steps — effectively turning your Completion criteria into something the runtime tracks automatically, with pause, resume, and edit controls, rather than something you have to restate each turn.
8. Enterprise rollout: AGENTS.md and shared templates
At team scale, retyping the same Context and Constraints into every prompt doesn't hold up. The pattern that scales is pushing repeatable context into a version-controlled AGENTS.md at the repository root — build commands, test commands, naming conventions, and forbidden patterns that apply to every task in that repo — so individual prompts only need to state what's unique to that specific piece of work: the Goal and any task-specific Constraints or Completion criteria.
- Standardize an
AGENTS.mdtemplate across repos (build/test commands, coding conventions, protected paths). - Keep a short library of Goal/Context/Constraints/Completion snippets for your team's most common task types (endpoint changes, migrations, refactors).
- Review completion criteria in code review the same way you'd review a test plan — a vague completion line is a signal the PR description needs work, not just the prompt.
AGENTS.md once, then every engineer's prompts are just a few lines of Goal, task-specific Constraints, and Completion — instead of a wall of repeated boilerplate.
9. Common mistakes
- Writing a Goal with a process verb instead of an outcome. "Refactor this code" and "make this better" don't define an end state; Codex (and you) can't tell when they're done.
- Treating Constraints as optional. Skipping them is how an agent that's technically allowed to edit files ends up adding a dependency, touching a schema, or editing a file nobody asked about.
- Writing Completion criteria that describe a feeling, not a check. "Looks clean" isn't verifiable; "the 11th request in a minute returns a 429" is.
- Repeating the same Context paragraph in every prompt. If you're pasting the same background every time, it belongs in
AGENTS.md, not in the prompt. - Bundling unrelated goals into one task. "Fix authentication, redesign pricing, and migrate the database" is three goals wearing one prompt; split them so each run has one definition of done.
❓ FAQ
A: For anything trivial and unambiguous, no. For anything where "done" isn't obvious — most real feature work and refactors — yes, especially Constraints and Completion criteria, which are the two most commonly skipped.
A: Context in a prompt is task-specific — "follow the pattern in this file." AGENTS.md is repo-wide and persistent — build commands, conventions, and protected paths that apply to every task, so you stop repeating them.
A: Goal mode makes the completion criteria persistent and runtime-tracked across many steps, with pause/resume/edit controls, instead of something only checked at the end of a single-turn prompt.
A: A runnable check is ideal, but it can also be a concrete manual verification step — the key requirement is that it's checkable, not that it's automated.
A: Use a planning step first — ask Codex to propose a plan or interview you — before locking in a Goal and Completion criteria, rather than starting implementation on a fuzzy objective.
- Prompting — OpenAI Developers
- Workflows — OpenAI Developers
- AGENTS.md — OpenAI Developers
- Using Goals in Codex — OpenAI Cookbook
Codex and OpenAI are trademarks of OpenAI. This post explains and synthesizes publicly available documentation and community practice in its own words; it does not reproduce any source verbatim. Feature availability (such as Goal mode) may change — verify against the live docs before relying on it in production.
📝 Summary
- Codex runs in a loop and needs a clear exit condition, or it invents its own.
- Goal — one concrete sentence, outcome verbs, no process verbs.
- Context — language, framework, and one existing pattern to mirror.
- Constraints — dependencies, files, and schemas that are off-limits.
- Completion criteria — a concrete, ideally runnable check.
- One worked example inside the prompt beats a paragraph of explanation.
- Goal mode extends the same idea to longer, multi-step work with persistent tracking.
- Push repeated Context into
AGENTS.mdso prompts stay short at team scale. - The most common mistake is a vague Goal or Completion line dressed up as detail.
Four short lines, one unambiguous target — that's the whole anatomy. Structure the prompt once, and Codex spends its effort building instead of guessing. 🚀
Comments
Post a Comment