In a production agent harness, permissions, action limits, human approval, and cancellation form the control layer between a model's proposed action and a real-world side effect.
That distinction becomes important the moment an agent can do more than produce text. Reading an order, changing an order, issuing a refund, deleting a resource, sending a message, or modifying a production system all have different consequences. A good harness does not give the model one giant "permission" switch. It turns authority into narrow, inspectable decisions.
This article develops that idea through one fictional support-operations agent. The scenario is invented for teaching. The architecture patterns are general, while vendor-specific capabilities are clearly identified where they are used as examples.
Think about a child using a family toolbox. Being able to see the screwdriver does not mean the child is allowed to use every tool, open every box, or repair the electrical panel. The harness is the set of rules that decides which tool can be used, on what, for how much, and whether an adult must approve it first.
OWASP's current GenAI security guidance describes excessive agency in terms of excessive functionality, excessive permissions, or excessive autonomy. It also recommends enforcing authorization at the downstream boundary instead of relying on an AI system to decide for itself whether an action is allowed. OWASP GenAI Security Project
Imagine a support agent handling an order problem. It can look up an order and calculate a refund recommendation automatically. Cancelling the order, however, changes customer-visible state. A large refund may also have financial consequences. The agent may therefore be allowed to propose those actions, while the harness decides whether they are permitted and whether a person must approve them.
02. Model choice is not authorization
03. Designing action limits
04. Human approval without approval fatigue
05. Cancellation: stop the run, stop the action, or undo the effect?
06. Build a durable control loop
07. The complete running example
08. Failure modes and Safety Checks
09. Enterprise Rollout
10. Common Mistakes
11. ❓ FAQ
12. 🔗 References & Further Reading
13. 📝 Summary
| Control | Primary question | Typical enforcement point | What it cannot guarantee |
|---|---|---|---|
| Permission | May this action be attempted? | Harness policy and downstream authorization | That the action itself is correct or safe in every context |
| Action limit | How much, how often, where, or how broadly? | Policy gate, quotas, scopes, budgets | That the remaining allowed action is desirable |
| Human approval | Should a person explicitly permit this consequential step? | Before the side-effecting operation | That the reviewer sees every hidden failure or future action |
| Cancellation | Should ongoing work stop? | Run controller plus cancellable tools | Undoing side effects that already happened |
01
What permissions mean inside an agent harness
A permission is not simply a property attached to a model. In an agent architecture, the model may propose a tool call, but another component should decide whether that call is acceptable for the current user, target, state, and policy.
This is why it helps to separate five responsibilities:
| Model | Interprets the task and may propose the next action. |
| Agent | Coordinates reasoning, tool selection, state transitions, and completion. |
| Harness | Controls runtime behavior: available tools, policy checks, approval gates, budgets, cancellation, state, tracing, and recovery. |
| Tool | Executes a defined operation against a real system. |
| Downstream system | Applies its own authentication, authorization, validation, and transaction rules. |
The exact boundaries vary. In one architecture, the application owns the whole runner. In another, an external platform may manage the agent loop and state. For example, OpenAI currently documents different agent runtime options where an Agents SDK application owns deployment, storage, approvals, and runtime integration, while its managed Agents API handles more of the harness infrastructure. That is a documented architecture distinction, not a universal definition of what every agent harness must look like. OpenAI Agents overview
The model can recommend an action; the harness should determine whether that action can cross the next trust boundary.
A tool appearing in the model's tool list does not have to mean every invocation is allowed. A useful harness can expose a capability while still constraining the target, arguments, amount, time window, identity, or approval state for each actual call.
Risk: Treating a tool's mere availability as permission can give an agent broader authority than the task requires.
Control: Evaluate authorization at the tool boundary and enforce it again in the downstream system.
Remaining risk: A correctly authorized action can still be wrong because the input, business state, or model interpretation is wrong.
Your agent can cross from information retrieval into an external side effect: financial operations, record changes, communications, deployments, deletions, access changes, or other consequential actions.
02
Model choice is not authorization
One of the easiest mistakes to make is assuming that a sufficiently capable model can be trusted to decide when an action is allowed. That turns a policy question into a prediction problem.
Consider two statements:
Question A: "Would cancelling this order help solve the customer's problem?"
Question B: "Is this caller authorized to cancel this order under today's policy?"
The first is a reasoning task. The second is an authorization task. The harness should not quietly turn the second into "ask the model what it thinks."
The distinction also matters for instruction injection. Information retrieved from documents, websites, tickets, emails, or tool outputs can contain text that attempts to influence the agent's next step. Even a useful model can misinterpret that text. The stronger boundary is to make authorization independent of the content that is merely being analyzed.
OWASP's excessive-agency guidance specifically recommends implementing authorization in downstream systems and applying the complete-mediation principle so requests made through tools are validated against security policy. OWASP LLM06:2025
The agent reads an order and determines that cancellation is a plausible remedy.
The harness then asks: Is the order owned by the current customer? Is the order still cancellable? Is the action inside the user's role? Is this a single-order action or a bulk action? Is approval required?
Only after those checks pass does the tool receive an executable request.
model_proposes(action)
↓
normalize_and_validate(action)
↓
check_identity_and_scope(action)
↓
check_action_limits(action)
↓
check_approval_policy(action)
↓
execute_only_if_allowed(action)
Policy should be a gate around the model, not a hope inside the prompt.
An agent handles untrusted external content, multiple users, sensitive tools, or actions whose authorization can be represented as explicit policy.
03
Designing action limits
A permission answers whether an action class is allowed. An action limit answers how far the system may go.
For an agent, "can cancel orders" is usually too broad. Production policy is more useful when it can express constraints such as:
| Limit type | Example | Why it matters |
|---|---|---|
| Capability | Read orders but do not delete records | Reduces unnecessary functionality |
| Resource scope | Only the current user's order IDs | Limits cross-user access |
| Quantity | One order cancellation per approved request | Prevents accidental bulk impact |
| Value | Refund above a defined business threshold requires review | Controls financial exposure |
| Destination | Only approved production tenant or account | Prevents unintended environments or recipients |
| Time | Approval expires after a short window | Avoids executing a stale decision |
| Concurrency | At most one destructive operation per customer session | Limits cascading side effects |
| Retry budget | Bound automatic retries on side-effecting calls | Prevents a transient failure from becoming repeated actions |
These limits are not meant to be decorative policy. They should be evaluated from trusted request metadata and authoritative system state. For example, the harness should not trust the model to tell it that an order belongs to the current user. The customer identity and order ownership should come from authenticated application state or the downstream service.
A library card may let you borrow books, but that does not mean you can borrow 500 books, use someone else's card, keep them forever, or remove books from the building. The limit is part of the permission.
A useful policy model is therefore not simply allow = true/false. Conceptually, it resembles:
identity,
action,
target,
current_state,
limits,
approval_state,
time_context
)
The exact implementation can be simple or elaborate. A small internal assistant might need only an allowlist and one approval gate. A regulated or high-impact system may need identity-bound scopes, transaction limits, operator controls, audit records, and independent downstream authorization.
Risk: A read-only-looking workflow accidentally receives a tool identity with write or delete privileges.
Control: Give the tool the minimum downstream authority required and enforce user-specific or task-specific scope in the destination system.
Remaining risk: A compromised or misbehaving agent can still misuse any capability it legitimately receives, so functionality reduction and action limits remain important.
The cost of a mistake grows with quantity, value, breadth, or repetition. Action limits are especially useful where "one bad call" and "one hundred bad calls" have very different consequences.
04
Human approval without approval fatigue
Human approval is most useful when a person is making a genuine authorization decision, not when the system simply asks for a click before every tool call.
The design question is therefore: which actions should pause?
A practical classification is:
| Action class | Typical treatment | Example reasoning |
|---|---|---|
| Read-only | Usually automatic | Reading an order status normally does not change external state. |
| Low-impact write | Policy-controlled | May be automatic if the effect is bounded, expected, and reversible. |
| Consequential write | Approval often appropriate | Cancellation, financial changes, publishing, access changes, or production mutations may need an explicit decision. |
| High-impact or irreversible action | Strong gate or separate workflow | The cost of an incorrect action is high enough that policy should require stronger evidence or human intervention. |
OpenAI's current Agents SDK documentation describes human review as a mechanism that pauses a run before sensitive tool work, while the decision itself is resolved through a resumable run state. It also recommends attaching checks close to the tool that creates the side effect rather than assuming that agent-level checks cover every custom tool in every orchestration pattern. OpenAI: Guardrails and human review
A teacher does not need to approve every pencil mark a student makes. The teacher should step in when the student is about to do something that can have a meaningful consequence. The goal is not maximum supervision. The goal is the right supervision at the right moment.
A strong approval request contains enough context to make the decision meaningful. At minimum, consider:
| Actor | Which authenticated user or service identity is requesting the action? |
| Exact action | What will actually happen if approved? |
| Target | Which order, account, environment, recipient, or resource will be affected? |
| Scope and magnitude | How many resources or how much value is affected? |
| Reason | Why does the agent believe this action is required? |
| Evidence | What authoritative information supports the proposed action? |
| Expiry | How long should the approval remain valid? |
Avoid a vague approval message such as "The agent wants to perform a sensitive operation. Approve?" A reviewer cannot meaningfully evaluate an action they cannot see.
Requested action: Cancel order <ORDER_ID>.
Customer: authenticated customer <CUSTOMER_ID>.
Current state: order is pending and remains cancellable.
Effect: order status changes from pending to cancelled.
Why proposed: customer requested cancellation before fulfillment.
Approval: explicit authorization required by policy.
This structure also supports auditing. Later, an operator can answer not just "Was approval given?" but "What exact action was presented, what information was visible, who approved it, and what actually executed?"
Risk: People approve repetitive prompts without inspecting them.
Control: Route only genuinely consequential actions to human review, show the exact proposed effect, and make approvals narrow and time-bounded.
Remaining risk: A reviewer can approve a bad action. Human review reduces one class of failure; it does not make the workflow infallible.
The action is consequential enough that a human decision adds meaningful control, but frequent enough that an explicit approval boundary can still fit the workflow.
05
Cancellation: stop the run, stop the action, or undo the effect?
"Cancel" sounds simple until the agent is already doing work.
Suppose the agent has completed three read operations, requested approval for a cancellation, and then started an external transaction. A user pressing Cancel can mean several different things:
| Cancellation layer | Meaning | Typical limitation |
|---|---|---|
| Stop scheduling | Do not begin another model or tool step. | Already-started work may continue. |
| Interrupt in-flight work | Ask the currently running tool or executor to stop. | Only possible if the execution layer supports safe cancellation. |
| Cancel pending work | Invalidate queued or approved-but-not-started actions. | Requires every action to check cancellation status before execution. |
| Compensate | Perform a separate recovery action after a side effect already happened. | Compensation is a new action, not time travel. |
This distinction is visible in current agent tooling. OpenAI's Agents API documents explicit cancellation of an active turn and explains that the session and previous work remain available. Its event reference also notes that cancellation can be accepted even after backend execution has ended, while previously saved results, published files, and existing terminal outcomes remain preserved. OpenAI: Run and continue sessions and OpenAI Agents session event reference
MCP provides another useful example. Its current Tasks extension defines explicit task cancellation, but describes cancellation as cooperative: a cancellation request signals intent, and the task can still end in a terminal state other than cancelled. That is an important systems lesson: a cancel signal is not automatically a proof that external work stopped. MCP Tasks extension
Pressing the stop button on a blender is not the same as putting the fruit back on the kitchen counter. Stop prevents more work, but it cannot undo material that has already been changed.
That leads to a critical rule:
Never document "cancel" as if it means "undo." Define exactly which state transition cancellation controls.
For a side-effecting tool, the harness should ideally know the state of the action:
↘
CANCELLING → CANCELLED
COMPLETED → COMPENSATION_REQUIRED → COMPENSATED
The actual state machine will differ by system. The important part is that the harness can distinguish an action that was merely proposed from one that already crossed the external side-effect boundary.
Risk: The UI says "Cancelled" while an external operation is still running or has already completed.
Control: Make cancellation state reflect confirmed runtime state, and distinguish requested cancellation from confirmed cancellation.
Remaining risk: Some downstream systems cannot stop an operation once committed; the recovery path then becomes reconciliation or compensation.
Tasks may run for a meaningful duration, call external services, consume compute, or create side effects that a user may reasonably need to interrupt.
06
Build a durable control loop
The strongest permission system is not a collection of conditionals scattered across tool handlers. It is a repeatable control loop.
One practical pattern is:
2. Establish caller identity and policy context
3. Let the agent reason and propose an action
4. Validate the requested tool and arguments
5. Re-check authoritative resource state
6. Apply action limits
7. Determine whether approval is required
8. Record a durable decision point
9. Execute the side effect
10. Verify the resulting state
11. Record outcome and evidence
12. Continue, stop, retry safely, or compensate
Step 8 deserves special attention. If approval is required, do not simply keep a boolean called approved=true. Associate the decision with the specific action that was reviewed.
A useful action record can contain an action identifier, authenticated actor, tool name, target resource, normalized arguments, a policy version, approval status, approval timestamp, expiry, execution status, and correlation information.
This gives you a stronger invariant:
Approval should authorize an exact action, not a vague future intention.
For example, "approved: cancel order" is weaker than "approved: cancel order <ORDER_ID> under policy version <POLICY_VERSION> before <EXPIRY_TIME>." The latter can be checked before execution.
The agent requests approval to cancel order A. A reviewer approves it. Before execution, the order changes to "shipped."
A weak implementation sees approved=true and proceeds.
A stronger implementation re-checks authoritative state, sees that the business condition changed, and blocks the action even though approval was previously granted.
This is one of the reasons approval should not replace authorization. Approval is a human decision in time; authorization is a policy decision against current state.
The same principle applies to cancellation. Before starting an approved action, the executor can check whether the run has been cancelled. For long-running operations, the tool or worker should also expose a cooperative cancellation path where the underlying system permits it.
action = agent.propose_tool_call()
assert policy.allows_tool(action.tool)
assert policy.allows_target(action.target)
assert policy.within_limits(action)
assert downstream_state.is_still_valid(action)
if policy.requires_approval(action):
approval = wait_for_explicit_decision(action)
assert approval.matches(action)
assert approval.is_unexpired()
assert run.is_not_cancelled()
result = execute(action)
verify(result)
record(action, result)
That last verify step matters. A tool returning an HTTP success, database acknowledgement, or generic "accepted" response is not necessarily the same thing as proving that the business effect occurred as intended.
Risk: A previously approved action executes after the world has changed.
Control: Re-authorize against current trusted state immediately before execution.
Remaining risk: Race conditions can still exist between authorization and commit, so the downstream system should use its own transaction and concurrency controls where appropriate.
Actions depend on state that can change between proposal, approval, and execution.
07
The complete running example
Let's connect the pieces with the fictional support-operations agent introduced earlier.
A customer says: "My order has not shipped. Please cancel it and refund me."
The agent has four capabilities:
1. Read order status.
2. Calculate refund eligibility.
3. Cancel an eligible order.
4. Submit a refund transaction.
Step 1 — identity is established. The application already knows which authenticated customer is making the request. The model does not get to choose the customer identity.
Step 2 — read-only inspection. The agent retrieves the order status and learns that the order is currently pending.
Step 3 — proposal. Based on the customer request and current information, the agent proposes cancellation and a refund.
Step 4 — authorization checks. The harness confirms that the order belongs to the customer context, cancellation is permitted in the current state, and the refund path is available.
Step 5 — action limits. The harness confirms that the request covers one order, one customer context, and a permitted amount.
Step 6 — approval decision. Suppose business policy requires human review for refunds above a threshold. Cancellation might be automatic while the refund waits for approval.
Step 7 — durable pause. The agent run pauses. The pending approval references the exact refund action rather than simply "refund customer."
Step 8 — the human approves. The system stores the approval decision and the relevant action identity.
Step 9 — re-check state. The harness verifies that the order is still eligible. If the order shipped while waiting, the action is rejected despite the earlier approval.
Step 10 — execution. The tool performs the approved operation using the appropriate downstream identity and scope.
Step 11 — verification. The harness reads the resulting order or transaction state and records evidence of the outcome.
Step 12 — cancellation during execution. If the customer presses cancel while a long-running downstream operation is still active, the harness requests cancellation. If the external system cannot stop the operation, the system must report that accurately and enter the appropriate recovery or reconciliation state.
The model interpreted the request, examined information, and proposed actions.
The harness bound the action to identity, policy, scope, approval, cancellation state, and execution lifecycle.
The tool translated the authorized request into an external system operation.
User
↓
Application / identity context
↓
Agent model
↓ proposes
Harness policy + limits + approval + cancellation
↓ authorizes
Tool adapter
↓
Downstream API / database / service
↓
Result + verification + trace
This is the main design lesson of the article: permission is a runtime decision about a concrete action, not a general promise that the agent is trustworthy.
Risk: The model is allowed to define the target, identity, or authorization context itself.
Control: Derive security-relevant context from trusted application state and validate tool arguments independently.
Remaining risk: The model can still make a poor recommendation inside its allowed scope. Quality evaluation and business validation are still needed.
You need a practical way to explain agent permissions to engineers, architects, reviewers, or security teams. A running scenario makes the separation of responsibilities easier to reason about than isolated definitions.
08
Failure modes and Safety Checks
Permission failures usually appear as ordinary engineering mistakes rather than dramatic security events. The useful question is not only "Can this happen?" but "What does the system do when it happens?"
| Failure | Consequence | Useful correction |
|---|---|---|
| Broad tool identity | A valid tool call reaches more data or write authority than intended. | Reduce downstream privileges and enforce user or task scope. |
| Approval detached from action | An approval can accidentally authorize a different action. | Bind approval to a concrete action record and expire it. |
| No state re-check | A stale decision executes against changed state. | Re-authorize immediately before execution. |
| Cancellation only at UI level | The interface says stopped while backend work continues. | Propagate cancellation into the runtime and tool layer where supported. |
| Retries repeat side effects | A transient error produces duplicate external actions. | Use idempotency where supported, narrow retry rules, and inspect execution state before retrying. |
| Approval fatigue | Reviewers begin approving without examining the action. | Reserve approval for meaningful decision points and present precise evidence. |
| Unverified completion | The agent reports success based on a tool response that is not proof of business completion. | Verify authoritative post-action state before declaring success. |
A cancellation path should not create a second uncontrolled channel that bypasses authorization, tracing, or cleanup. The system should know which run or action is being cancelled, who requested it, and what state the cancellation request reached.
Instruction injection deserves special attention because it can influence an agent's decision without directly changing the harness policy. For example, a retrieved document could contain text attempting to persuade the agent to call an unrelated administrative tool. The harness should not treat that instruction as an authority grant. Tool policy, identity, and downstream authorization must remain independent.
No prompt filter, sandbox, approval step, or single validation layer completely eliminates this class of risk. The safer design is defense in depth: limit available capabilities, constrain permissions, validate action arguments, require approval for appropriate high-impact operations, isolate execution where needed, and make consequential activity observable.
Risk: Untrusted content causes the agent to request an action outside the user's intended task.
Control: Treat external content as data, keep authorization outside model reasoning, enforce least privilege, and require approval for appropriate side effects.
Remaining risk: The agent may still choose a harmful action that falls within its legitimate authority, so business rules and post-action verification remain necessary.
You are moving from a demo agent to a production workflow and need to test not only successful paths, but also stale approvals, retries, cancellation races, unauthorized targets, and partial execution.
09
Enterprise Rollout
Enterprise adoption is less about adding a single approval screen and more about assigning ownership for the control surface.
Security ownership: Define which controls must be enforced by identity systems, gateway policy, the harness, tools, and downstream applications.
Application ownership: Define what the agent is allowed to do, what requires review, and what evidence proves completion.
Operations ownership: Define who can stop runs, revoke access, disable tools, or activate an emergency deny policy.
Audit ownership: Define retention, traceability, approval evidence, and access to consequential action records.
Change ownership: Version action policies and review changes as carefully as other production control logic.
NIST's current work on software and AI agent identity and authorization explicitly focuses on identifying, managing, and authorizing access and actions taken by software agents. Its 2026 concept paper also frames increased agent autonomy as a reason for stronger identity and authorization controls. NIST: Software and AI Agent Identity and Authorization
Phase 1: Start with read-only tools and establish identity, traceability, and completion checks.
Phase 2: Add bounded low-risk writes with explicit action limits.
Phase 3: Introduce human approval at clearly defined high-impact boundaries.
Phase 4: Add cancellation propagation, durable action state, recovery, and compensation paths where required.
Phase 5: Test the whole workflow under stale state, retries, partial failures, denied approvals, and malicious external content.
Do not assume that every architecture needs every component. A simple internal agent may not need a full sandbox or multi-stage approval workflow. Conversely, an agent performing production infrastructure changes may need substantially stronger isolation and authorization than a support chatbot.
The control surface should be proportional to the consequence of failure.
Do not copy a security architecture because another agent platform uses it. Start with your own trust boundaries, side effects, identities, and failure costs.
Risk: The organization treats vendor defaults as its enterprise authorization model.
Control: Map the agent workflow to enterprise identities, downstream permissions, incident controls, and explicit ownership before production rollout.
Remaining risk: Governance controls can become stale. Policies and access must be reviewed as tools, workflows, and business rules change.
You are turning an agent prototype into an enterprise service with multiple teams, production identities, incident procedures, compliance requirements, or customer-impacting side effects.
10
Common Mistakes
Mistake 1 — "The model has the tool, so it must be allowed."
Why it happens: developers treat tool registration as authorization. Correction: separate capability exposure from per-action authorization and enforce downstream permissions.
Mistake 2 — "Approve this category forever."
Why it happens: broad approvals are convenient. Correction: bind approvals to concrete actions, scoped identities, and expiry conditions.
Mistake 3 — "Cancel means undo."
Why it happens: the UI hides execution state. Correction: distinguish cancellation requests, execution stopping, and compensation.
Mistake 4 — "The approval is valid because the user clicked it."
Why it happens: approval is treated as a final authorization token. Correction: re-check live business state immediately before the side effect.
Mistake 5 — "Retry the tool until it works."
Why it happens: retry logic is copied from read-only integrations. Correction: side-effecting operations require stricter retry semantics, state inspection, and idempotency mechanisms where supported.
Mistake 6 — "Ask for human approval on everything."
Why it happens: humans are seen as the safest generic control. Correction: reserve approval for meaningful decision points and use automatic policy checks for routine low-risk actions.
Mistake 7 — "The tool succeeded, so the task succeeded."
Why it happens: transport success is confused with business success. Correction: define completion criteria and verify authoritative state before reporting completion.
Mistake 8 — "Permissions are only a security-team concern."
Why it happens: authorization is separated from application design. Correction: developers, architects, security, operations, and business owners should agree on the action policy before release.
When reviewing an agent tool, ask five questions: Who? What? Where? How much? What happens if we need to stop? Those five questions often expose missing controls faster than reading a large tool description.
Risk: The team tests only happy-path prompts.
Control: Test denial, stale approvals, changed state, cancellation races, duplicate retries, unauthorized targets, tool outages, and partial side effects.
Remaining risk: Testing samples behavior. Production systems can still encounter conditions the test suite does not cover.
11
❓ FAQ
No. Approval is most useful for meaningful consequential actions, while low-risk operations can often be controlled through explicit policy, narrow permissions, and action limits.
No. Authorization determines whether an action is permitted under policy, while approval is an explicit human decision that may be required for a particular action at a particular time.
No. Cancellation can stop future work or interrupt cooperative operations, but an external side effect that already committed may require separate verification or compensation.
Because the world can change between approval and execution, making the previously approved action stale or invalid.
The harness should enforce its own action policy, but the downstream system should also enforce authorization so a model or tool cannot bypass the real security boundary.
12
🔗 References & Further Reading
The article uses the following primary sources to verify current architecture and security claims. The explanations, examples, analogies, tables, and pseudocode are original to this article.
OpenAI — Guardrails and human review
Human review, approval interruptions, tool-level controls, and approval-state handling.
OpenAI — Running agents
Agent loop, continuation, and runtime behavior.
OpenAI — Agents overview
Current distinctions among managed and application-controlled agent runtime options.
OpenAI — Run and continue sessions
Current agent-session cancellation behavior.
OpenAI — Agent session event reference
Cancellation event semantics and lifecycle details.
OWASP GenAI Security Project — LLM06:2025 Excessive Agency
Excessive functionality, excessive permissions, excessive autonomy, downstream authorization, and human approval.
NIST — Software and AI Agent Identity and Authorization
Current identity and authorization work for software and AI agents.
Model Context Protocol — Tasks extension
Long-running task state, input requirements, and cooperative cancellation.
Model Context Protocol — 2026-07-28 specification release
Current protocol changes relevant to stateless operation, multi-round interaction, authorization, and tasks.
Vendor names, product names, specifications, and API identifiers belong to their respective owners.
13
📝 Summary
1. An agent can propose an action without being authorized to execute it.
2. Permissions should be narrow and enforced at the tool and downstream security boundaries.
3. Action limits control scope, quantity, value, destination, time, concurrency, and retries.
4. Human approval is most useful at consequential decision points, not as a blanket confirmation screen.
5. Approval should bind to a concrete action and should not survive a meaningful change in authoritative state without another check.
6. Cancellation means different things at different layers; stopping a run does not automatically undo external side effects.
7. Production-grade harnesses need durable action state, verification, tracing, controlled retries, and recovery paths.
8. No single safeguard makes an agent completely safe; the goal is bounded authority and controlled failure.
Comments
Post a Comment