Skip to main content

AI Governance: Securing AI Applications, Agents, and Memory

Calculating read time…
Securing an AI application is not the same thing as securing the model inside it. 

A deployed AI system can expose data through retrieval, inherit malicious instructions from documents, call business tools with more authority than intended, preserve poisoned information in memory, or allow a single mistaken decision to travel across several automated steps. The governance question is therefore larger than “Is the model secure?” It is: what can this AI system access, influence, remember, and do, under whose authority, with which controls, and with what evidence?

Consider a fictional internal assistant called Natkhat. At first, Natkhat only answers questions from company documents. Later, the business adds customer-record lookup, draft-message creation, and a narrowly scoped record-update capability. The model has not suddenly become a different kind of technology. The system has become more consequential. A malicious document can now influence an action. A stolen session can matter more. A poisoned memory entry can affect later conversations. A weak permission boundary can turn a misleading response into a real-world change.

This post treats security as a governance problem: identify the trust boundaries, constrain authority, test the controls, retain evidence, monitor changes, and explicitly accept or reduce the remaining risk. The examples are fictional and are intended for teaching, not as descriptions of any real organization.

01
What “Secure” Means in AI Governance

Security, safety, governance, assurance, and legal compliance often appear in the same discussion, but they answer different questions. A good AI governance process keeps those questions connected without pretending they are interchangeable.

🟧 Child-friendly analogy

Think about a school laboratory. Security is the locked door, controlled keys, safe storage, and protection from tampering. Safety asks whether the experiment itself could hurt someone. Governance decides who is allowed to run the experiment, who approves higher-risk work, what records must be kept, and what happens when something goes wrong. A laboratory can have excellent locks and still run a dangerous experiment. An AI system can be technically hardened and still have poor governance.

Area Core question Typical evidence What it does not prove
Security Can unauthorized people, content, or components gain influence or access? Access tests, attack simulations, logs, configuration reviews That the system is safe or legally compliant in every context
Safety Could system behavior create unacceptable harm, even without an attacker? Safety evaluations, incident analysis, human review That confidentiality and access controls are adequate
Governance Who decides, approves, operates, monitors, and accepts residual risk? Policies, ownership records, approvals, control evidence That a control works merely because it is documented
Assurance What evidence supports confidence in the system or control? Independent testing, repeatable results, audit trails A guarantee that future behavior cannot fail
Compliance What binding requirements apply to this organization and use case? Legal analysis, regulatory records, required controls That a voluntary framework is itself law

For this article, the central governance unit is therefore not the model in isolation. It is the deployed AI system and its surrounding control environment: user identity, application logic, model provider, retrieval sources, memory stores, tools, credentials, agents, operators, affected people, monitoring, and escalation paths.

Governance principle

A model response is information. A system action is an authority decision. The governance burden rises when generated content can cross a trust boundary and change something outside the model.

The NIST AI Risk Management Framework is voluntary and is intended to help organizations manage AI risks across the lifecycle. NIST also states that AI RMF 1.0 is being revised, so this article does not treat the framework as a frozen or mandatory security standard.

🎯 Use this when...

A project team says, “The model provider handles security.” Ask instead: which layer is the provider securing, which layers remain under our control, and what evidence do we have for the complete deployed system?

02
Map the Real AI Security Boundary

Security reviews often start with the model API. That is usually too narrow. An AI application can be exposed through everything that supplies context, stores state, authorizes actions, or interprets generated output.

A practical governance inventory should answer at least six questions: who calls the system, what data can enter it, what state persists, what tools it can reach, what identity those tools see, and what happens after the model responds.

User / operator
  ↓
Application authentication + session policy
  ↓
Request handling + policy enforcement
  ↓
Model / agent reasoning
  ↙                ↘
Trusted business data      Untrusted external content
  ↓                   ↓
Retrieval / memory gateway ← validation / provenance / isolation
  ↓
Tool proposal → policy check → authorization → tool execution
  ↓
Business system / communication channel
  ↓
Audit trail + monitoring + incident response

The diagram is a governance model, not a prescribed architecture. The important principle is the placement of control points. A model should not become the sole enforcement point for permissions merely because it can describe what it intends to do.

🟧 Tricky concept: “The model asked for it” is not authorization

Imagine an employee asks a receptionist to open a restricted room. The receptionist should not decide, based only on the employee’s persuasive explanation, that the employee must therefore have access. An AI tool call deserves the same separation: the model can propose an operation, while an authorization layer decides whether that operation is actually permitted.

A useful boundary map records responsibility as well as technology. The AI provider may control model infrastructure. The deployer may control the application, data, policies, and user access. A tool vendor may control an external service. An operator may control incident response. The affected person may have no technical control at all, even though a system action can affect them.

Governance Check

Risk → The team assesses only the model endpoint and misses retrieval stores, memory, tools, service identities, or external data sources.

Control → Maintain a system-level inventory showing data flows, persistent state, connected services, identities, action capabilities, owners, and trust boundaries.

Evidence → Current architecture record, system inventory entry, dependency list, access matrix, data-flow record, and named control owners.

Remaining risk → Third-party components and undocumented data flows may still change faster than the inventory. Reconciliation must therefore be periodic, not one-time.

NIST’s Generative AI Profile is useful here because it frames risk management around the AI lifecycle rather than treating a model as the whole system. Its scope is cross-sectoral and voluntary.

🎯 Use this when...

You are reviewing a vendor diagram. Redraw the system from the perspective of what can influence the AI and what the AI can influence. Those two directions usually reveal missing controls.

03
Identity, Permissions, and Tool Authority

The largest security change in an agent often occurs when the system moves from answering to acting. A text response can be wrong. A privileged action can make the wrong answer operational.

This is why least privilege matters especially strongly for agents. Give the system the smallest set of capabilities needed for the approved task. Separate read access from write access. Separate drafting from sending. Separate reversible operations from irreversible ones. Use independent authorization checks where the impact warrants them.

🟧 Child-friendly analogy

Give a child a library card before giving them the master key to the building. A library card might allow borrowing books. A master key can open offices, storage rooms, and equipment areas. The technical equivalent is capability scope: an agent that only needs to read customer order status should not inherit permission to delete customers or send external messages.

A useful way to govern agent authority is to define an action ladder. The precise categories will vary by organization, but a simple model is:

Action class Example Default control Evidence
Read Retrieve an approved record Identity + data-level authorization Access log
Draft Prepare a message but do not send Output validation + user review Draft + review record
Reversible write Change a low-impact preference Policy + transaction check Before/after audit record
High-impact / irreversible Release funds or send an external commitment Explicit authorization and meaningful human approval where required by the risk decision Approval, transaction record, identity, timestamp

This ladder is an original governance recommendation, not a universal control standard. The organization must set thresholds based on the consequence of failure. A low-risk internal lookup may not need the same approval path as a transaction affecting a customer, employee, or financial position.

Worked Example — Natkhat’s authority boundary

Natkhat may read an approved support record and prepare a response. It may not independently send external email. It may propose a customer-record correction, but the application must validate the target record, the field being changed, the allowed value range, and the user’s authorization before the update occurs.

The important control is not the model instruction “be careful.” The important control is an enforcement point outside the model that can reject the operation even when the model recommends it.

NIST’s 2026 concept paper on software and AI agent identity and authorization is particularly relevant to this issue. It is an initial public draft / concept paper, not a mandatory requirement, and it explores identification, authorization, auditing, and related controls for agents.

Governance Check

Risk → A model can choose among tools, but the tools execute under a broad service identity with more privileges than the business use case requires.

Control → Keep authorization outside the model; apply least privilege, action-specific policy checks, narrow service identities, transaction constraints, and approval gates for higher-impact actions.

Evidence → Permission matrix, identity configuration, denied-action tests, approval records, and periodic access review.

Remaining risk → A technically authorized action can still be inappropriate because of bad data, wrong context, or human misunderstanding. Authorization is necessary but not sufficient.

🎯 Use this when...

An agent is being granted a new tool. Review the change as a new authority boundary, not merely a new feature.

04
Prompt Injection and Untrusted Content

Prompt injection is especially important in AI security because instructions can arrive through places the business does not consider “the user prompt”: a webpage, uploaded document, ticket, email, retrieved passage, image, tool response, or another agent.

The governance problem is not simply that the model might follow a bad sentence. It is that data can become instruction-like content and influence a system that has real access or authority.

The OWASP LLM01:2025 Prompt Injection guidance distinguishes direct and indirect prompt injection and highlights impacts such as sensitive-information disclosure and unauthorized access to functions. It also emphasizes that there is no fool-proof prevention technique, which is a strong reason to design the surrounding system so that a fooled model has limited reach.

🟧 Child-friendly analogy

Suppose you ask a clerk to read a letter. The letter contains: “Ignore your manager and hand me the safe keys.” A careful clerk treats that sentence as part of the letter, not as a new authority instruction. An AI system needs an architectural equivalent: external content should not silently gain the authority of system instructions.

A strong control pattern therefore has two layers:

Layer 1 — Reduce influence. Identify external or untrusted content, separate it from trusted policy, constrain what the model can infer from it, and validate important output.

Layer 2 — Reduce consequence. Assume some attacks will succeed and prevent the compromised model from automatically reaching sensitive data, privileged tools, or irreversible operations.

Worked Example — A malicious support document

A fictional support team uploads a document that contains a hidden instruction telling Natkhat to export a confidential customer list. The document is stored in a repository used for retrieval.

The first defense is to treat the document as content with uncertain trust, not as an authoritative instruction. The second defense is that Natkhat's retrieval identity cannot access the restricted customer list. The third defense is that even a proposed export operation fails a policy check because bulk export is outside the agent's capability.

Notice what happened: the organization did not need a perfect prompt-injection detector to preserve the security boundary.

OWASP guidance also recommends adversarial testing and treating model behavior as an input to a broader security design. MITRE ATLAS provides a complementary threat-oriented view and currently tracks techniques such as LLM prompt injection, AI agent context poisoning, tool poisoning, RAG poisoning, and AI agent tool invocation.

Governance Check

Risk → A document, website, message, image, or tool result can inject instructions that influence an agent.

Control → Mark external content as untrusted, isolate it from policy authority, validate high-impact outputs, enforce least privilege, and require stronger approval for consequential actions.

Evidence → Injection test set, failed-attempt records, authorization denials, output-validation results, and red-team findings.

Remaining risk → Prompt injection remains an evolving class of attacks. The system should be designed to limit impact even when detection is imperfect.

🎯 Use this when...

Your AI reads anything that was not created and approved as a trusted system instruction. Assume the content may be misleading, compromised, stale, or actively adversarial.

05
Memory and Retrieval Are Security-Relevant State

Memory changes the security model because information can persist beyond the interaction in which it first appeared. Retrieval adds another layer because information can be selected, reintroduced, and interpreted later.

A useful governance distinction is:

State Why it matters Governance questions
Conversation context Influences the current response Who can see it? How long does it persist? What happens when the session changes?
Long-term memory Can affect future behavior Who can write it, inspect it, correct it, expire it, or delete it?
Retrieval store Determines which information is reintroduced Are access rules preserved during retrieval? Can poisoned or stale entries enter the context?
Agent state / configuration May influence future plans or actions Can state be changed without review? Is provenance known? Is configuration treated as trusted control data?

Security controls for memory should therefore include more than encryption at rest. Depending on the use case, a governance design may need provenance, write authorization, source classification, data minimization, tenant isolation, retention rules, expiration, correction and deletion paths, integrity monitoring, and testing for cross-user or cross-session leakage.

OWASP's May 2026 discussion of memory as an attack surface highlights the security significance of persistent context. The broader governance lesson is useful even without adopting any particular product architecture: persistent AI state should not automatically be treated as trustworthy simply because the system stored it.

🟧 Tricky concept: storage is not trust

A sealed box can contain either a trusted document or a poisoned document. The fact that the box is inside your building does not answer whether the contents deserve authority. Memory requires both storage controls and trust decisions.

Retrieval systems also create authorization questions. An embedding store can contain information that is logically restricted even when its representation is not human-readable. OWASP's guidance on vector and embedding weaknesses discusses risks including unauthorized access and leakage when access controls or retrieval design are inadequate.

Governance Check

Risk → A malicious or incorrect piece of information becomes persistent memory and influences future sessions or users.

Control → Restrict who and what can write memory, attach provenance, classify trust, apply retention and expiry, preserve authorization boundaries during retrieval, and provide correction or deletion mechanisms where appropriate.

Evidence → Memory schema, write-path controls, provenance records, retention configuration, isolation tests, deletion tests, and sampled retrieval audits.

Remaining risk → A legitimate memory can become misleading when circumstances change. Periodic review and expiry may therefore be necessary even when no attacker is present.

🎯 Use this when...

An agent can remember preferences, instructions, summaries, facts, prior decisions, or tool results beyond the current request. Treat the memory write path as a controlled security function.

06
Secrets, Data, and Trust Boundaries

AI systems often move sensitive data through more places than application teams initially expect: prompts, attachments, retrieval indexes, caches, logs, traces, memory, tool requests, model-provider APIs, and downstream applications. Security governance should therefore ask not only “Can the user access this data?” but also “Can this AI path access, transform, retain, or expose it?”

A mature design separates at least four ideas:

Data authorization: whether the requesting identity may access the underlying information.

Data minimization: whether the AI system needs all of the information it can technically receive.

Data handling: where the information is stored, logged, cached, retrieved, transmitted, and deleted.

Secret handling: whether credentials are held by controlled infrastructure rather than exposed to model-generated text or long-lived prompts.

🟧 Child-friendly analogy

Do not give a delivery driver your building's master key simply because they need to deliver one package. The same principle applies to AI tools. A model may need the ability to request a transaction, but it should not receive a reusable secret that grants unrestricted access to the underlying system.

Illustrative policy pattern — not a legally sufficient policy and not vendor-specific code

MODEL → proposes operation only
POLICY GATE → validates identity, target, scope, and action type
CREDENTIAL BROKER → supplies controlled authorization to the tool
TOOL → executes only the approved operation
AUDIT LOG → records actor, decision, target, result, and timestamp
MODEL ← receives constrained result, not reusable secret material

The important architectural distinction is between instruction and capability. Instructions describe what the AI should do. Capabilities determine what the surrounding system can actually do. Security should not depend entirely on an instruction remaining effective after an adversarial input reaches the model.

Worked Example — Avoiding a credential-shaped permission

Natkhat needs to update a controlled field in a business application. Instead of placing a broad API credential in the prompt or model context, the application exposes a narrow operation through a policy-enforced gateway. The gateway checks the authenticated user, permitted record scope, allowed field, and transaction rule before obtaining the necessary backend authorization.

If the model is manipulated, the gateway still makes the final authorization decision.

NIST's SP 800-218A is relevant to secure development of AI models and AI-enabled software. It is a community profile that supplements the SSDF and is not a replacement for the broader operational security architecture. Its own scope notes are important: the final profile focuses on AI model development and does not cover deployment and operation of AI systems in the same way an end-to-end security program must.

Governance Check

Risk → Sensitive data or reusable credentials appear in prompts, memory, logs, traces, or model-accessible context.

Control → Minimize data, segregate secrets, restrict retention, mask or transform sensitive values where appropriate, and keep privileged authorization in controlled application components.

Evidence → Data-flow assessment, log review, secret-scanning results, retention settings, access tests, and configuration evidence.

Remaining risk → Even with strong controls, downstream systems may create additional copies or exposure paths. Vendor and dependency reviews remain necessary.

07
Testing, Monitoring, and Change Control

A security control that is never tested is a claim about design, not evidence about operation. AI systems make this especially important because model behavior, prompts, retrieval data, tools, policies, and surrounding software can all change.

A proportionate security evaluation should cover at least four kinds of tests:

Boundary tests: Does the application reject access outside the user's authorized scope?

Adversarial tests: What happens when untrusted content tries to influence the model or tools?

Abuse-path tests: Can a sequence of individually permitted actions combine into a harmful outcome?

Recovery tests: Can the team detect, contain, investigate, and recover from a poisoned memory, compromised integration, or incorrect privileged action?

Illustrative security evaluation record

Control ID: AI-SEC-017
Risk: Prompt manipulation could influence a privileged tool request.
Control: Tool gateway independently validates identity, scope, and action type.
Test: 40 approved adversarial cases + 10 negative authorization cases.
Expected result: No unauthorized tool execution.
Evidence: Test run ID, denial logs, reviewer sign-off, configuration snapshot.
Decision owner: AI system owner + security risk owner.
Residual risk: Novel attack paths may evade the current test set.
Next review trigger: Model change, tool addition, policy change, or material incident.

The record above is completely illustrative. It is not a required NIST form, an OWASP certification artifact, or a universal audit template.

Monitoring should also observe the security consequences of AI behavior rather than only technical availability. Depending on the use case, useful signals may include unusual tool-call frequency, denied authorization attempts, unexpected data access, sudden changes in memory writes, repeated requests around control boundaries, abnormal destinations, and high-impact actions that lack an expected approval trail.

🟧 Tricky concept: retesting is not “starting over”

A small change can move a security boundary. Adding one new tool, changing the memory schema, replacing a retrieval source, modifying authorization logic, or switching model behavior can invalidate an earlier assurance argument. The right response is targeted reassessment, not automatically a complete audit of everything—or blindly trusting the old evidence.

Governance principle

Change management should be risk-triggered, not merely calendar-driven. A new model, new tool, new memory behavior, new data source, or new user population can change the risk profile before the next scheduled review.

The current OWASP GenAI LLM Top 10 2026 is described by OWASP as its latest community-driven guide for critical security risks in LLM applications. Separately, the OWASP Agentic AI work focuses on risks that become more significant when systems can plan, coordinate, use tools, and act. These are useful inputs to a security testing program; they are not substitutes for application-specific threat modeling or authorization design.

🎯 Use this when...

A team says, “We tested the AI last quarter.” Ask what changed since that test and which changes could create a new attack path.

08
Implementation in Practice

The goal of implementation is not to create a giant checklist. It is to create a chain from risk → policy → technical control → test → evidence → residual risk → owner.

Step 1 — Describe the system in business language. Write down what the AI is for, who uses it, what decisions or actions it supports, which people can be affected, and what happens when it fails.

Step 2 — Enumerate capability. List data sources, retrieval stores, memory, tools, identities, external integrations, agent-to-agent communication, and administrative interfaces.

Step 3 — Classify action impact. Separate read, draft, reversible, and high-impact actions. Decide which actions require independent authorization or human approval.

Step 4 — Mark trust boundaries. Identify where user input, third-party content, retrieved data, memory, model output, tool response, or another agent can enter the workflow.

Step 5 — Put enforcement outside the model where necessary. Use application controls for authorization, data scope, transaction constraints, and approval gates. Do not ask a language model to be the only enforcement layer for a privileged operation.

Step 6 — Test both failure and attack paths. Include realistic misuse, not only happy-path accuracy tests. Test denial, escalation, persistence, cross-user boundaries, tool misuse, and recovery.

Step 7 — Retain evidence. Keep the smallest useful set of records that makes the security decision reviewable: approved design, permission boundaries, tests, findings, approvals, exceptions, monitoring evidence, and reassessment triggers.

Step 8 — Assign residual risk to a real owner. Someone with appropriate authority must accept, reduce, transfer, or reject the remaining risk. “The security team reviewed it” is not the same as “the organization accepted the remaining business risk.”

Worked Example — Natkhat security decision
Business need Help authorized support users find information and prepare customer responses.
Material risk Untrusted content could influence the agent to access or change data outside the intended scope.
Control Identity-aware retrieval, narrowly scoped tools, external authorization, high-impact approval, memory write controls, and audit logging.
Test Injection attempts, unauthorized-record requests, cross-user retrieval tests, malicious-memory scenarios, and tool-authorization denials.
Evidence Architecture record, access matrix, test results, denied-action logs, approval records, and exception register.
Residual risk Novel prompt-injection techniques, unexpected interactions among trusted components, and operator error remain possible.
Decision Illustrative approval for limited deployment, subject to the stated controls and reassessment triggers.

All details above are fictional. The example demonstrates the structure of a governance decision; it does not represent a certification criterion.

For larger environments, this process can be connected to an AI inventory, enterprise identity management, secure software development, vendor risk management, incident response, and internal audit. The point is integration: AI security should not become an isolated spreadsheet maintained by a team that cannot enforce the resulting controls.

Governance Check

Risk → Security documentation exists, but no one owns the decision to accept the remaining risk.

Control → Record named decision rights, control owners, reviewers, escalation paths, approval thresholds, and reassessment triggers.

Evidence → Signed or attributable decision record, control ownership register, exception record, and review history.

Remaining risk → Ownership itself can become stale after organizational changes. Revalidate accountability when system scope or organizational responsibilities change.

09
Common Mistakes

1. “The prompt says not to do that.” A system prompt is an important instruction mechanism, but it is not equivalent to authorization enforcement. Correction: place critical permission checks in deterministic application controls.

2. “The vendor has security covered.” A provider can secure its service while the deployer still misconfigures retrieval, identity, memory, tools, logging, or data access. Correction: separate provider controls from system-level controls and gather evidence for the complete boundary.

3. “Human approval means the action is safe.” A user who sees a polished AI recommendation may approve it without understanding the hidden context or attack path. Correction: make approval meaningful by showing the material action, target, impact, relevant evidence, and opportunity to reject or modify it.

4. “Memory is just a convenience feature.” Persistent state can influence future behavior and therefore deserves access, provenance, retention, and integrity controls. Correction: govern memory writes and retrieval boundaries as security-relevant operations.

5. “We passed the initial security test.” A new tool, data source, model, memory behavior, or authorization change can create a new risk path. Correction: define change-triggered reassessment criteria before production deployment.

6. “We can fix prompt injection with better filtering.” Detection can reduce exposure, but no single filter should carry the entire security argument. Correction: combine input handling, trust separation, least privilege, output validation, authorization, and consequence limitation.

7. “Security is the security team’s responsibility.” Security teams can review controls, but business ownership, system ownership, data ownership, engineering ownership, and risk acceptance may sit elsewhere. Correction: define decision rights across the lifecycle.

Governance principle

The strongest AI security design assumes the model may be fooled and asks how little damage a fooled model can cause.

10
❓ FAQ

Q1. Is prompt injection mainly a model problem or an application security problem?

It is both, but the governance impact depends heavily on the surrounding application. Prompt injection becomes substantially more consequential when the AI can access sensitive data, persistent memory, privileged tools, or external systems. A sound design therefore addresses model behavior and limits what a manipulated model can reach.

Q2. Should every AI agent require human approval before every tool call?

No. Requiring a person to approve every low-risk operation can create fatigue without improving security proportionately. The better approach is risk-based: apply stronger authorization and meaningful human approval where the action has material consequences, while using automated controls for lower-risk operations.

Q3. Does encrypted memory solve the security problem for AI memory?

No. Encryption can protect stored data from some forms of unauthorized access, but it does not answer whether a memory entry is trustworthy, whether the correct user can retrieve it, whether it should still exist, or whether poisoned information can influence future behavior. Memory security requires both protection of the storage and governance of the information inside it.

Q4. What evidence should an AI governance reviewer ask for before approving privileged agent actions?

At minimum, the reviewer should be able to trace the approved capability to a named owner, the identity and permission boundary to the intended use case, the control to a specific enforcement point, the control to test evidence, and the remaining risk to an accountable decision-maker. For higher-impact actions, the review should also show how approval, monitoring, and incident response work in practice.

Q5. Can a well-designed AI security program guarantee that an agent will never be compromised?

No. Security controls reduce likelihood, limit impact, improve detection, and support recovery; they do not create certainty about every future attack, model failure, dependency failure, or human mistake. A credible governance program therefore documents residual risk and reassesses it when the system, threat environment, or business impact changes.

11
🔗 References & Further Reading

NIST AI Risk Management Framework — Voluntary AI risk-management framework; NIST notes that AI RMF 1.0 is being revised.

NIST AI 600-1 — Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile — Cross-sectoral profile for managing generative-AI risks.

NIST AI 100-2 E2025 — Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations — Terminology and taxonomy for AI security threats and mitigations.

NIST SP 800-218A — Secure Software Development Practices for Generative AI and Dual-Use Foundation Models — AI-focused supplement to the SSDF.

NIST 2026 Initial Public Draft — Software and AI Agent Identity and Authorization — Draft concept paper exploring agent identity, authorization, auditing, and related controls; it is not a binding standard.

OWASP GenAI LLM Top 10 2026 — Current OWASP guidance for major security risks in LLM applications.

OWASP LLM01:2025 Prompt Injection — Direct and indirect prompt-injection risks and mitigation approaches.

OWASP LLM08:2025 Vector and Embedding Weaknesses — Security risks around retrieval, embeddings, access control, and leakage.

OWASP Agentic AI — Threats and Mitigations — Threat-model-oriented guidance for emerging agentic risks.

OWASP Top 10 for Agentic Applications — Agent-specific risk categories including identity and privilege abuse, memory and context poisoning, and insecure inter-agent communication.

OWASP Agent Control Standard (ACS) — September 2026 resource focused on inspectability, traceability, instrumentation, and runtime control of agents.

OWASP — Memory Is a Feature. It Is Also an Attack Surface — Discussion of persistent memory and context as security-relevant state.

MITRE ATLAS — Living knowledge base of adversary tactics and techniques involving AI systems.

Framework names, project names, and other protected names belong to their respective owners.

12
📝 Summary

• Secure the system, not only the model. Include identity, data, retrieval, memory, tools, agents, operators, and affected people.

• Treat model output as a proposal when consequences matter. Independent authorization should decide what the system may actually do.

• Assume untrusted content can influence the model. Prompt injection defenses should limit both the chance of manipulation and the damage a manipulated system can cause.

• Treat memory as security-relevant state. Govern who can write it, what is trusted, how long it persists, who can retrieve it, and how it can be corrected.

• Convert governance statements into enforceable controls. “Require approval” is incomplete without a trigger, owner, enforcement point, test, evidence, and residual-risk decision.

• Reassess after material change. New tools, models, memory behavior, data sources, users, and integrations can change the security boundary.

Final takeaway

A secure AI system is not one that never fails. It is one whose authority is deliberately bounded, whose failures are observable, whose controls are tested, and whose remaining risk is consciously owned.


Comments