Skip to main content

Governing Third-Party Models, Tools, and Vendors

Calculating read time…
Most organizations do not build every AI component they use. They may obtain a foundation model from one provider, a document-processing service from another, a retrieval component from a third party, an external tool connector from another supplier, and supporting infrastructure from yet another provider. An AI application can therefore become a chain of dependencies rather than a single product.

That creates a governance problem that is easy to underestimate: How do you remain accountable for an AI system when important parts of that system are outside your organization?

A vendor questionnaire alone does not answer that question. Neither does a security certificate, a model card, a contract, an impressive benchmark, or a statement that a provider follows responsible-AI practices. Those materials may provide useful evidence, but governance has to connect them to the actual system you deploy, the people affected by it, the actions it can take, and the risks your organization is willing to accept.

Consider a fictional internal assistant called Northstar Assistant. It initially answers questions from company documents using a third-party language model. Later, the organization gives it two additional tools: one to retrieve customer records and another to send approved notifications. The model, document service, hosting platform, identity provider, and tool connectors may all come from different suppliers.

The governance question is no longer simply “Is the model trustworthy?” It becomes “Can we justify, operate, monitor, control, and if necessary disable the complete system built from these dependencies?”

Governance principle

A supplier can own a component. Your organization still needs an explicit owner for the decision to use that component within your AI system.

📑 In This Post

01. Why third-party AI changes the governance problem

02. Know exactly what you are buying and integrating

03. Perform risk-based vendor and dependency due diligence

04. Turn procurement requirements into enforceable controls

05. Validate the vendor component inside your own system

06. Govern third-party tools when AI can take actions

07. Monitor change, incidents, concentration, and exit risk

08. Implementation in practice: the Northstar example

09. Common mistakes and how to correct them

10. FAQ

References. Primary-source reading

Summary. The practical governance test

🔀 Quick Comparison
Dependency What you need to understand Typical governance concern Useful evidence
Third-party model Capabilities, limitations, versions, interfaces, usage conditions, data handling Output behavior, confidentiality, change risk, service dependency Provider documentation, test results, approved version, contract terms
Third-party tool Functions, permissions, identities, inputs, outputs, downstream systems Unauthorized or unintended actions Permission design, test evidence, logs, owner approval
AI service vendor Service scope, support model, subprocessors, availability, changes Operational dependency and service changes Contract, service commitments, incident process, change notices
Open-source component Maintainer activity, dependencies, license, release history, security posture Supply-chain integrity and unsupported dependencies Version inventory, vulnerability review, license review, provenance records

01
Why Third-Party AI Changes the Governance Problem

A useful starting point is to stop thinking about procurement as something that happens before AI governance. For AI systems, the dependency relationship itself becomes part of the governance surface.

🍊 Child-friendly analogy

Imagine hiring a school bus company to transport children. You did not build the bus, hire the driver, or maintain the road. But you still need to know where the bus goes, who is allowed to drive it, what happens if the bus breaks down, and what you will do if the company suddenly changes the route.

AI governance works similarly. The external provider may control an important component, but the deploying organization controls how that component is selected, configured, connected, permissioned, and used within the business process.

This is particularly important because the risk of the complete application can be different from the risk of the model or tool in isolation. A language model used only to summarize public documents is not governed in the same way as the same model connected to internal payroll records and a message-sending tool.

Governance principle

Assess the system you are deploying, not just the component you are buying.

This distinction also prevents a common governance error: transferring responsibility rhetorically to the vendor. A supplier may have responsibility for its own service, but statements such as “the provider handles AI safety” are incomplete unless the organization can show exactly what the provider controls, what the customer controls, and how the two sets of controls connect.

The NIST AI RMF Playbook is useful here because its current material explicitly includes third-party AI systems, third-party data, procurement, supply-chain concerns, testing, auditability, and contingency planning. The Playbook is voluntary, and NIST currently notes that AI RMF 1.0 is being revised. That makes it a governance reference rather than a universal legal requirement. NIST AI RMF Playbook

🛡 Governance Check

Risk → The organization evaluates a vendor component but never evaluates the complete AI workflow.

Control → Require system-level risk assessment after the external component is integrated into the intended workflow.

Evidence → Architecture record, approved dependency inventory, integration test results, risk decision, named system owner.

Remaining risk → Vendor behavior and external changes can still create new conditions after approval.

🎯 Use this when...

Your AI system depends on models, data, tools, hosted services, external APIs, managed infrastructure, open-source packages, or other components that your team does not fully control.

02
Know Exactly What You Are Buying and Integrating

Before asking whether a vendor is trustworthy, define what the vendor actually provides. “AI vendor” is too broad to be a useful governance category.

A governance inventory should distinguish at least the following:

Dependency question Why it matters Example evidence
Who provides the component? Clarifies accountable commercial and technical relationships. Supplier identity, contract, service identifier
What exact component is used? A vendor may offer multiple models, versions, regions, APIs, or tool packages. Model or component identifier and approved version
What data crosses the boundary? Different data categories create different privacy, confidentiality, and contractual concerns. Data-flow map, classification, retention setting
What can the component do? Functions and permissions determine potential impact. API specification, tool manifest, permission matrix
What can change without your code changing? Hosted AI services can change models, behavior, policies, limits, or supporting infrastructure. Change policy, version pinning, notification terms

The inventory should also record where the dependency sits in the system. A provider might supply the model itself, a retrieval service, a moderation service, an embedding service, an agent framework, a tool, a data source, or a complete managed application.

This distinction is important because each dependency may create a different control boundary. For example, a model provider may control model behavior, while your organization controls authorization. A tool provider may expose a business action, while your organization controls whether that action is available to the AI agent at all.

🟢 Worked Example — Dependency Map

The fictional Northstar Assistant has five external dependencies:

  1. A hosted language model for generation.
  2. A document-processing service for extracting text from uploaded files.
  3. A customer-data API used by a retrieval tool.
  4. A notification service used to send approved messages.
  5. Open-source orchestration libraries used by the application.

The governance inventory treats each as a separate dependency. The assistant is then evaluated as a system that combines them.

🎯 Use this when...

A team says “we use Vendor X for AI” but cannot clearly state which model, service, tool, version, data path, identity, and downstream system are actually involved.

03
Perform Risk-Based Vendor and Dependency Due Diligence

Not every supplier requires the same depth of review. A model that summarizes public marketing material presents a different governance problem from a model that assists with a high-impact decision or controls access to internal records.

A practical approach is to classify the dependency according to the consequence of failure, not merely the sophistication of the technology.

Illustrative tier Example use Review depth Decision
Lower impact Drafting or summarizing non-sensitive information Basic supplier review, data checks, functional testing Business owner approval
Material impact Internal records, employee support, customer communication Expanded privacy, security, legal, technical and operational review Named risk owner and documented decision
High impact Consequential decisions or AI-enabled actions with significant downstream effect Deep technical evaluation, contractual controls, scenario testing, monitoring and contingency Formal risk acceptance or escalation according to organizational policy

A good due-diligence review should answer several different questions rather than collapsing everything into one supplier score.

Capability: What does the component actually do, and what does it explicitly not guarantee?

Data: What information is processed, stored, retained, logged, or potentially used for service improvement, subject to the applicable contract and service configuration?

Security: How is access controlled? What identities, credentials, network paths, isolation boundaries, and security processes are relevant?

AI-specific behavior: What evaluation information, limitations, known failure modes, versioning, and usage restrictions are documented?

Operational resilience: What happens during an outage, degradation, provider incident, rate-limit event, or significant service change?

Legal and rights: What rights, licensing conditions, intellectual-property concerns, data-processing obligations, or jurisdictional conditions could affect the intended use?

Exit: Can the organization replace the component, disable it, recover needed information, and continue the business process safely?

🍊 Tricky concept — Evidence is not the same as assurance

A vendor may provide a security certificate, audit report, model card, policy, benchmark, or compliance statement. That is evidence about something the supplier says, measures, or has had assessed. It does not automatically prove that your specific configuration, data, workflow, or AI-assisted decision is safe or appropriate.

This distinction is especially important for management-system standards. ISO/IEC 42001 specifies requirements for an AI management system, while ISO/IEC 23894 provides guidance on AI risk management. Neither should be treated as a universal declaration that a particular third-party model, tool, or deployment is safe for every use case. See the official ISO/IEC 42001 and ISO/IEC 23894 references.

🛡 Governance Check

Risk → Procurement accepts a supplier based primarily on reputation or generic certifications.

Control → Require evidence to be mapped to the actual dependency, use case, data, permissions, and risk tier.

Evidence → Completed due-diligence record, evidence register, technical test results, contract review, risk decision.

Remaining risk → Evidence can become outdated and may not reveal failures unique to the organization's integration.

04
Turn Procurement Requirements into Enforceable Controls

One of the weakest patterns in AI governance is to write a strong-sounding policy statement and assume the risk has been controlled.

Policy statement: “Sensitive data must not be exposed to unapproved AI providers.”

Weak control interpretation: “The team was told not to do it.”

Stronger control interpretation: “Only approved AI endpoints are reachable through the production integration path; data classification rules block prohibited fields; service identity is restricted; exceptions require documented approval; test cases verify the block.”

The second formulation is stronger because it connects a policy to an actual technical or procedural mechanism and makes it possible to collect evidence.

For third-party AI, contractual language and technical controls should reinforce one another. A contract cannot compensate for an architecture that sends data to the wrong endpoint. A technical control cannot resolve every legal or commercial question. Governance is stronger when both layers agree.

Governance requirement Possible control Evidence
Use only approved providers Approved-provider allowlist and controlled integration credentials Current approved inventory and configuration review
Protect sensitive data Classification rules, filtering, minimization, access controls Test results, configuration records, data-flow review
Control provider changes Change notification requirement and internal re-evaluation trigger Change log, review ticket, retest results
Handle incidents Incident notification path and internal response procedure Incident record, communication log, corrective action
Enable safe exit Replacement or shutdown procedure with dependency inventory Exit test or documented recovery procedure

The exact contractual terms depend on the service, jurisdiction, bargaining position, data sensitivity, and organizational policy. Governance teams should not assume that a generic AI addendum is suitable for every use case.

The same principle applies to legal requirements. Different laws can allocate duties according to the role played by an organization, the type of AI system, the geography, and the use case. The European Union AI Act is a useful illustration: Article 25 sets out circumstances in which certain distributors, importers, deployers, or other third parties can be treated as providers of a high-risk AI system, including specified cases involving branding or substantial modification. That is a legal rule for the applicable circumstances, not a universal vendor-governance principle for every jurisdiction.

The AI Act also places specific obligations on providers of general-purpose AI models. As of 2026, those obligations are operational matters for providers in scope; organizations integrating such models still need to understand their own role and obligations rather than assuming that the upstream provider has solved every downstream governance issue.

Useful primary sources are the EU AI Act text and the European Commission's current GPAI obligations guidance.

🎯 Use this when...

Your governance program contains policies and contracts but reviewers struggle to demonstrate how those statements are actually enforced in the running system.

05
Validate the Vendor Component Inside Your Own System

Vendor due diligence answers what the supplier provides and what evidence the supplier can present. It does not answer whether the component behaves acceptably in your environment.

A third-party model may perform well on a provider's evaluations and still be unsuitable for a specific workflow because your documents, language mix, instructions, tool chain, user population, or downstream business rules are different.

The same is true of tools. A tool can be technically secure yet be inappropriate for the AI workflow because its permissions are broader than the task requires.

🟢 Worked Example — From Vendor Evidence to System Evidence

Suppose a provider supplies an AI model with documentation describing intended use and known limitations. Northstar Assistant should not simply attach that document to its governance record and mark the dependency approved.

Instead, the team could test representative internal questions, deliberately ambiguous requests, prohibited requests, incorrect premises, sensitive-data cases, and situations in which the model's answer is used by a downstream business process.

The result becomes application-specific evidence. It demonstrates what the complete workflow does under the organization's chosen configuration rather than what the model documentation alone claims.

Evaluation should be connected to the risk. There is little value in collecting dozens of generic model metrics if the material business risk concerns unauthorized retrieval, wrong customer identification, inappropriate message sending, or failure to stop after ambiguous input.

A practical evaluation plan can include:

  1. Representative normal-use scenarios.
  2. Known failure and edge cases.
  3. Security and misuse cases appropriate to the architecture.
  4. Permission-boundary tests for connected tools.
  5. Human-review and escalation tests.
  6. Regression tests after model, tool, prompt, policy, or integration changes.

The organization's evidence should also show what version or configuration was tested. Otherwise, months later, it can become impossible to establish whether a failed outcome occurred under the same dependency state that was originally approved.

🛡 Governance Check

Risk → Procurement evidence is treated as if it were evidence that the integrated AI system works safely.

Control → Require application-specific evaluation before production approval and after material changes.

Evidence → Test plan, test inputs, observed results, pass/fail criteria, dependency version, reviewer and approval date.

Remaining risk → Testing covers selected scenarios, not every possible input or future operating condition.

🎯 Use this when...

A vendor's documentation looks convincing but the team cannot demonstrate how the dependency behaves when exposed to the organization's actual data, users, tools, and business workflow.

06
Govern Third-Party Tools When AI Can Take Actions

Third-party tools deserve special attention because they can convert model output into external effects.

A model that only generates text has one risk surface. A model that can call a tool to update a record, create an order, send a message, execute code, or modify a document has a much larger operational surface.

This is where the distinction between model governance and agent or application governance becomes critical. The model may be supplied by Vendor A, while the tool comes from Vendor B and the target business system is controlled by your organization. The behavior emerges from the combination.

🍊 Child-friendly analogy

Giving a child a calculator does not give the child permission to change the school's bank account. The important question is not only whether the calculator works. It is what the child is allowed to do with the result.

For AI tools, that means reviewing at least functionality, permissions, identity, scope, confirmation, and reversibility.

Control question Better governance answer
Does the tool expose more functions than required? Remove or disable unnecessary functions.
Does the tool use excessive permissions? Use the narrowest practical identity and downstream permissions.
Can the AI trigger a high-impact action directly? Introduce confirmation, approval, policy enforcement, or another independent decision mechanism when warranted.
Is the action reversible? Prefer reversible operations or compensating controls for material actions.
Can the organization determine what happened? Maintain appropriate audit records for requests, tool calls, approvals, outcomes and failures.

This area aligns closely with current AI-security guidance. OWASP's 2025 LLM guidance describes risks from supply-chain vulnerabilities and excessive agency, while its material on excessive agency highlights excessive functionality, excessive permissions, and excessive autonomy as important causes of risk. Its agentic work separately identifies issues such as agentic supply-chain vulnerabilities and identity or privilege abuse. See the official OWASP LLM Top 10, Excessive Agency guidance, and OWASP Agentic Applications material.

A human approval step also needs careful design. A person clicking “Approve” after seeing only a vague AI-generated sentence is not necessarily meaningful oversight. A stronger control gives the reviewer enough information to understand what will happen, to whom, with what data, under which authority, and why.

Illustrative approval record — not a legally sufficient template

Action: Send customer notification
Requested by: Northstar Assistant workflow
Target: <CUSTOMER_ID>
Reason: Account-status update
Source data: Internal account record
Validation checks: Passed
Reviewer: <AUTHORIZED_REVIEWER>
Decision: Approved
Time: <TIMESTAMP>
Result: Notification service accepted request
Follow-up: Retain outcome for audit trail

Notice that the approval is attached to a defined action rather than simply to a conversation. That design makes the control easier to understand and test.

🛡 Governance Check

Risk → A third-party tool gives the AI broader downstream authority than the business task requires.

Control → Minimize tool functions, scope identities, separate read and write capabilities where practical, and introduce independent checks for higher-impact actions.

Evidence → Permission matrix, tool configuration, authorization tests, approval records, audit logs.

Remaining risk → Even narrowly scoped tools can be misused through unexpected inputs or workflow combinations.

07
Monitor Change, Incidents, Concentration, and Exit Risk

Approval is not the end of third-party governance. The dependency is still part of the production system after go-live.

A supplier might introduce a new model version, modify a service, change rate limits, alter a data-handling practice, add or replace subprocessors, experience an incident, discontinue an interface, or materially change pricing and availability.

Not every change should trigger the same response. The governance program should define which events require reassessment.

Change event Possible governance response
Model or major service version change Review scope and rerun relevant evaluation scenarios.
Material data-handling change Reassess privacy, contractual and architectural implications.
Security or AI incident Activate incident response and determine whether use should be restricted or suspended.
Service degradation or outage Use fallback, graceful degradation, manual process, or temporary shutdown as appropriate.
Supplier exit or product discontinuation Execute replacement or retirement plan.

There is another risk that deserves more attention as organizations standardize on a small number of powerful AI providers: concentration risk. If dozens of business processes depend on the same external AI service, a single provider outage, policy change, commercial decision, or significant technical change can affect many applications simultaneously.

Avoiding concentration risk does not automatically mean operating several providers. Multi-vendor architecture can add cost, complexity, data-transfer risk, inconsistent behavior, and additional governance work. The right decision depends on the business impact of failure and the realistic alternatives.

A sensible exit plan does not have to mean maintaining an immediately interchangeable second vendor. It can mean knowing exactly how to disable the dependency, switch to a manual process, preserve necessary records, communicate the interruption, and prevent the AI workflow from silently failing.

NIST's current Playbook specifically includes contingency considerations for high-risk third-party AI systems and suggests addressing third-party failures through organizational procedures and, where appropriate, redundancy. Again, these are voluntary risk-management suggestions rather than universal legal requirements.

🛡 Governance Check

Risk → A supplier changes or fails, but the organization discovers the impact only after users report a problem.

Control → Define material-change triggers, monitoring responsibilities, supplier notifications, incident procedures and fallback operations.

Evidence → Dependency monitoring record, change-review history, incident procedure, fallback test, reassessment decision.

Remaining risk → Some supplier changes or failures may occur before the organization receives reliable notice.

🎯 Use this when...

A production AI system depends on an external service whose behavior, availability, security posture, or commercial terms can change after deployment.

08
Implementation in Practice: The Northstar Example

The following is a fictional teaching scenario. It does not describe a real organization's controls, supplier selection, or production environment.

Northstar Assistant begins as an internal document-answering application. It uses a third-party language model and a document retrieval service. The initial risk is moderate because the system is intended to answer employee questions from approved internal material and does not directly modify business records.

Later, business leaders request two new capabilities: retrieve customer-account information and send customer notifications.

That change materially alters the governance problem. The organization should revisit the assessment rather than simply adding two buttons.

Decision stage Northstar decision Evidence retained
1. Define purpose Assistant supports internal information retrieval; customer communication is initially excluded. Purpose statement and scope boundary
2. Inventory dependencies Model, document service, hosting, identity and integration dependencies recorded. Dependency register and architecture record
3. Assess vendor risk Data handling, security, service dependency, change process and contract reviewed. Due-diligence record and evidence register
4. Validate system Representative internal scenarios tested against the configured application. Evaluation results and approval record
5. Add customer lookup New data path and access permissions assessed separately. Updated data-flow map and permission tests
6. Add notification tool Sending messages classified as an external action requiring additional controls. Tool assessment, approval mechanism and audit design

A possible fictional approval rule could be:

Illustrative governance rule — fictional, not legal advice

IF action = information retrieval
AND requested data = within approved user scope
AND authorization check = passed
THEN allow retrieval

IF action = external notification
THEN require policy validation
AND require authorized human approval
AND record the target, reason, content, approver and outcome

IF provider version changes materially
THEN suspend automatic promotion
AND trigger reassessment

Notice how the governance mechanism becomes more specific as the system's impact increases. The rule does not simply say “use AI responsibly.” It identifies the action, boundary, trigger, approval and evidence.

It is also important to identify who owns each decision. In this fictional example, the system owner might own the business purpose and operational decision; security might own relevant security controls; privacy or legal functions might advise on applicable obligations; the tool administrator might operate permissions; and a designated risk owner might accept or escalate residual risk according to internal policy.

Those roles do not automatically belong to the vendor.

🟢 Worked Example — A Compact Risk Record

Risk: A changed third-party model could alter customer-notification behavior.

Impact: Incorrect or inappropriate customer communication.

Control: Material model changes trigger regression evaluation; notification actions require an independent approval gate.

Owner: Fictional Northstar system risk owner.

Evidence: Version record, regression results, approval log and monitoring record.

Residual risk: Testing cannot predict every future model behavior or every novel user input.

🎯 Use this when...

A third-party component is moving from an informational role into a workflow that can affect records, communications, transactions, decisions, or other external outcomes.

09
Common Mistakes and How to Correct Them

Third-party AI governance often fails in predictable ways. The problem is usually not that an organization has no controls; it is that the controls are disconnected from the actual dependency and workflow.

Mistake Why it fails Better approach
“The vendor is responsible for AI safety.” The vendor does not control every downstream configuration, user, tool and business process. Map provider controls and customer controls separately, then evaluate the complete system.
“The contract handles it.” Contract language does not automatically enforce technical behavior. Map important contractual commitments to system controls and verification evidence.
“The certification proves the system is safe.” Certification or assessment generally addresses a defined scope rather than every deployment condition. Treat external assurance as one evidence source among several.
“The model passed our first test.” Model behavior, integrations and business context can change. Define retesting triggers and maintain regression suites.
“Human approval makes it safe.” A human cannot meaningfully review an action without sufficient context and authority. Design meaningful review with clear action, target, rationale, evidence and decision rights.
“One supplier review is enough.” Dependency behavior and business context can change after approval. Make third-party governance continuous and trigger-based.

Another frequent mistake is excessive paperwork without decision clarity. A 60-page questionnaire is not necessarily better governance than a concise review that identifies the actual data boundary, permission model, failure mode, owner, test, evidence and residual risk.

The aim is not to remove uncertainty. AI governance cannot eliminate uncertainty. The aim is to make material uncertainty visible enough that the organization can decide what to accept, what to control, what to monitor, and what to reject.

🛡 Governance Check

Risk → The organization accumulates documents but cannot explain why a supplier dependency was approved.

Control → Every material governance decision should identify the purpose, risk, control, evidence, owner and residual risk.

Evidence → Decision record linked to supporting documents and tests.

Remaining risk → Documentation quality depends on keeping the record current as the system evolves.

10
❓ FAQ

Q1. Does using a reputable AI provider remove the need for internal AI governance?

No. A reputable provider can reduce some supplier risk, but the organization still needs to assess how the service is configured, what data it receives, which users can access it, what actions it can trigger, and whether the resulting system is appropriate for the intended use.

Q2. Is a vendor questionnaire enough for third-party AI due diligence?

Usually not for material AI uses. A questionnaire can collect useful information, but higher-impact dependencies also require evidence review, system-specific testing, architectural analysis, permission assessment, contractual review and continuing monitoring proportionate to the risk.

Q3. Should every AI application use multiple model providers to avoid vendor lock-in?

No. Multiple providers can reduce some concentration risks but can also introduce cost, technical complexity, inconsistent behavior and additional governance work. The appropriate approach depends on business criticality, realistic alternatives, switching difficulty and the consequences of provider failure.

Q4. Does a human approval step make an AI-enabled action adequately governed?

Not automatically. The reviewer needs enough information, authority and time to make a meaningful decision. The control should define what is being approved, identify the target and consequence, enforce the decision, and retain appropriate evidence of the review.

Q5. What is the most important document to maintain for a third-party AI dependency?

There is no single universally sufficient document. A strong governance record usually connects the dependency inventory, intended use, data flows, owners, supplier evidence, controls, tests, contractual considerations, monitoring triggers, incident path, approval decision and residual risk.

11
🔗 References & Further Reading

Framework, standard, legal and project names belong to their respective owners.

12
📝 Summary

Third-party AI is still part of your AI system.

Vendor reputation is evidence, not a substitute for your own governance decision.

Inventory the actual model, tool, data path, identity, version and downstream dependency.

Make due diligence proportional to the consequences of failure.

Convert policies and contracts into enforceable controls and reviewable evidence.

Test the third-party component inside the complete application, not only in the vendor's context.

Be especially deliberate when third-party tools give AI systems the ability to act.

Monitor material changes, incidents, concentration risk and the ability to exit safely.

The strongest third-party AI governance record is not the one with the largest number of questionnaires. It is the one that lets another reviewer understand what was approved, why it was approved, who owns the decision, which safeguards are operating, what evidence supports them, and what risk remains.


Comments