Skip to main content

The Anatomy of an Effective Codex Prompt: Goal, Context, Constraints, Completion

Calculating read time…

An effective Codex CLI prompt is really a tiny spec, and reliably good ones share four ingredients: a Goal, Context, Constraints, and Completion criteria. Miss one, and Codex either guesses at what you meant, picks its own libraries and patterns, or stops halfway through because it isn't sure what "done" means. 🎯

Here's why this matters more than a style preference: Codex works in a loop, calling the model and then acting on file reads, edits, and commands until it decides the task is finished. If your prompt doesn't define completion, Codex has to invent its own finish line — and an invented finish line is exactly where "close enough" implementations, half-applied refactors, and silently skipped edge cases come from. A well-structured prompt spends your Codex usage on implementation instead of clarification. 🧩

Diagram showing the four elements of an effective Codex prompt: Goal, Context, Constraints, and Completion Criteria, flowing into agent execution

🔀 Quick Comparison — the four elements at a glance
Element Answers Missing it causes
Goal What outcome am I asking for? Vague, open-ended implementations
Context What should Codex look at first? Off-pattern code that ignores your codebase's conventions
Constraints What must Codex avoid? New dependencies, touched schemas, or edited files you didn't want changed
Completion criteria How does Codex know it's finished? Work that stops halfway, or continues past "good enough"

1. Why Codex needs this structure at all

Codex isn't a single-shot autocomplete — it's an agent running a loop of model calls and tool actions (file reads, edits, shell commands) until it judges the task complete. That loop needs an exit condition, and if you don't supply one, the model has to infer it. A prompt like "make this better" or "refactor this code" gives Codex no reliable finish line, so it either stops too early on a convenient proxy (like "the tests I happened to touch still pass") or keeps making changes you never asked for. The four-element structure exists to remove that ambiguity before the loop even starts.

2. Goal: what "done" looks like

The Goal is one concrete sentence describing the observable outcome, not the steps to get there. Use outcome verbs — add, fix, migrate, replace — rather than process verbs like improve or refactor, which describe an activity but not an end state.

Goal: Add per-API-key rate limiting to the /api/search endpoint.
💡 Contrasting example: "Make the search endpoint more robust" is not a Goal — there's no way for Codex, or you, to check whether it succeeded. "Add rate limiting" is checkable: either excessive requests get rejected or they don't.

3. Context: what Codex should look at first

Context points Codex at the language version, framework, and — critically — an existing pattern in your own codebase to mirror, rather than whatever pattern the model would default to on its own. Without this, Codex may quietly choose a different library, a different throttling strategy, or a different response format than the rest of your project uses, simply because nothing told it not to.

Context: Python FastAPI service, Redis already used for caching, follows
the token-bucket pattern already implemented in src/api/auth.py for
throttling failed login attempts.

Every time you find yourself copy-pasting the same background paragraph into prompt after prompt, that's a signal it belongs in a persistent AGENTS.md file at the repo root instead — more on that in the enterprise section below.

4. Constraints: what Codex must never do

Constraints are the boundaries: forbidden dependencies, protected files, schema changes that are off the table. This is the element people skip most often, and it's the one that causes the most expensive mistakes, because an agent that can edit files and run commands will happily "solve" a problem in a way you didn't want if nothing told it not to.

Constraints: Do not add new dependencies (redis client is already in
requirements.txt). Do not change the response schema for successful
requests. Do not modify the /api/auth endpoint.
✅ Worked example: a team asks Codex to add rate limiting and explicitly rules out new dependencies. Codex implements the token-bucket counters using the Redis client already in requirements.txt, instead of quietly pulling in a separate rate-limiting library nobody reviewed.

5. Completion criteria: the check that ends the loop

This is the exit condition for the agent's loop, and it should be something concrete and — ideally — runnable, not a subjective judgment call. "The code looks good" isn't verifiable. "Requests beyond the configured limit return HTTP 429, and pytest exits with code 0" is something both you and Codex can check.

Completion: A client sending more than 10 requests per minute to
/api/search receives a 429 response with a Retry-After header on the
11th request. Existing tests still pass. Run pytest.

🎯 Use this when: the task has any ambiguity at all about what "finished" means — which, in practice, is almost every non-trivial task.

6. The full worked example, line by line

Put together, the four elements read as a single, unambiguous prompt:

Goal: Add per-API-key rate limiting to the /api/search endpoint.

Context: Python FastAPI service, Redis already used for caching, follows
the token-bucket pattern already implemented in src/api/auth.py for
throttling failed login attempts.

Constraints: Do not add new dependencies (redis client is already in
requirements.txt). Do not change the response schema for successful
requests. Do not modify the /api/auth endpoint.

Completion: A client sending more than 10 requests per minute to
/api/search receives a 429 response with a Retry-After header on the
11th request. Existing tests still pass. Run pytest.

Notice what each line buys you: the Goal stops Codex from wandering into unrelated refactors; the Context stops it from inventing a new throttling strategy instead of reusing the token-bucket approach already proven in auth.py; the Constraints stop it from adding a dependency or touching an endpoint you didn't ask about; and the Completion criteria give it — and you — an objective way to know the run is over. One specific example like this is worth more in a prompt than a paragraph of general explanation.

7. From one-off prompts to Goal mode

The four-element structure works well for single-session tasks, but longer, multi-step work benefits from Codex's built-in Goal mode, started with /goal (enable it first with features.goals = true in config.toml if it isn't already on). A goal acts as both the starting prompt and the persistent completion criteria Codex checks against as it works across many steps — effectively turning your Completion criteria into something the runtime tracks automatically, with pause, resume, and edit controls, rather than something you have to restate each turn.

💡 When not to reach for Goal mode: if the finish line is genuinely vague — "make this better," with no defined end state — Goal mode won't rescue a poorly defined objective. Fix the Completion criteria first; a goal only tracks a target you've actually specified.

8. Enterprise rollout: AGENTS.md and shared templates

At team scale, retyping the same Context and Constraints into every prompt doesn't hold up. The pattern that scales is pushing repeatable context into a version-controlled AGENTS.md at the repository root — build commands, test commands, naming conventions, and forbidden patterns that apply to every task in that repo — so individual prompts only need to state what's unique to that specific piece of work: the Goal and any task-specific Constraints or Completion criteria.

  1. Standardize an AGENTS.md template across repos (build/test commands, coding conventions, protected paths).
  2. Keep a short library of Goal/Context/Constraints/Completion snippets for your team's most common task types (endpoint changes, migrations, refactors).
  3. Review completion criteria in code review the same way you'd review a test plan — a vague completion line is a signal the PR description needs work, not just the prompt.
✅ Rollout pattern that scales: a payment-services team keeps build/test commands and conventions in AGENTS.md once, then every engineer's prompts are just a few lines of Goal, task-specific Constraints, and Completion — instead of a wall of repeated boilerplate.

9. Common mistakes

  • Writing a Goal with a process verb instead of an outcome. "Refactor this code" and "make this better" don't define an end state; Codex (and you) can't tell when they're done.
  • Treating Constraints as optional. Skipping them is how an agent that's technically allowed to edit files ends up adding a dependency, touching a schema, or editing a file nobody asked about.
  • Writing Completion criteria that describe a feeling, not a check. "Looks clean" isn't verifiable; "the 11th request in a minute returns a 429" is.
  • Repeating the same Context paragraph in every prompt. If you're pasting the same background every time, it belongs in AGENTS.md, not in the prompt.
  • Bundling unrelated goals into one task. "Fix authentication, redesign pricing, and migrate the database" is three goals wearing one prompt; split them so each run has one definition of done.

❓ FAQ

Q: Do I need all four elements for every prompt?
A: For anything trivial and unambiguous, no. For anything where "done" isn't obvious — most real feature work and refactors — yes, especially Constraints and Completion criteria, which are the two most commonly skipped.
Q: What's the difference between Context in a prompt and AGENTS.md?
A: Context in a prompt is task-specific — "follow the pattern in this file." AGENTS.md is repo-wide and persistent — build commands, conventions, and protected paths that apply to every task, so you stop repeating them.
Q: How is Goal mode different from just writing good Completion criteria?
A: Goal mode makes the completion criteria persistent and runtime-tracked across many steps, with pause/resume/edit controls, instead of something only checked at the end of a single-turn prompt.
Q: Should Completion criteria always be an automated test?
A: A runnable check is ideal, but it can also be a concrete manual verification step — the key requirement is that it's checkable, not that it's automated.
Q: What if I don't know the Goal precisely up front?
A: Use a planning step first — ask Codex to propose a plan or interview you — before locking in a Goal and Completion criteria, rather than starting implementation on a fuzzy objective.
🔗 References & Further Reading

Codex and OpenAI are trademarks of OpenAI. This post explains and synthesizes publicly available documentation and community practice in its own words; it does not reproduce any source verbatim. Feature availability (such as Goal mode) may change — verify against the live docs before relying on it in production.

📝 Summary

  • Codex runs in a loop and needs a clear exit condition, or it invents its own.
  • Goal — one concrete sentence, outcome verbs, no process verbs.
  • Context — language, framework, and one existing pattern to mirror.
  • Constraints — dependencies, files, and schemas that are off-limits.
  • Completion criteria — a concrete, ideally runnable check.
  • One worked example inside the prompt beats a paragraph of explanation.
  • Goal mode extends the same idea to longer, multi-step work with persistent tracking.
  • Push repeated Context into AGENTS.md so prompts stay short at team scale.
  • The most common mistake is a vague Goal or Completion line dressed up as detail.

Four short lines, one unambiguous target — that's the whole anatomy. Structure the prompt once, and Codex spends its effort building instead of guessing. 🚀

Comments