Vendor AI Governance and Due Diligence: Managing Third-Party AI Risk
Vendor AI governance and due diligence puts EU AI Act liability on the deployer. Learn what evidence to demand, who signs, and how to start small.
When a vendor’s model screens out a job candidate or declines a loan, who answers for it: the company that built the model, or the organisation that deployed it? Under the EU AI Act, the deployer answers. Vendor AI Governance and Due Diligence is the discipline that closes that gap before a regulator, or a rejected applicant, finds it.
What Is Vendor AI Governance and Why It Matters for Third-Party Risk
Vendor AI governance is the deployer’s own accountability layer for AI systems built by someone else: the documented evidence, decision rights and monitoring that let a deploying organisation show it used a third-party system lawfully and safely.
A vendor management programme that a board risk committee has already approved usually stops short of that layer, and the shortfall sits exactly where the organisation is most exposed. The programme was designed to answer whether a supplier is available, secure and delivering what the contract describes. An AI system can satisfy all three and still produce a discriminatory decision on a Tuesday afternoon.
Deployer Liability Under the EU AI Act
The EU AI Act makes the deploying organisation directly answerable for how it uses a high-risk AI system, whoever built it. Article 26 sets the core deployer duties: use the system in line with the provider’s instructions, assign human oversight to people with the competence and authority to exercise it, make sure the input data is relevant to the system’s intended purpose, monitor operation, keep the logs the system generates, and tell the provider and the authorities when something goes wrong. Article 27 adds a fundamental-rights impact assessment for certain deployers, including those in credit scoring and in public services.
These duties attach to the deployer on its own account and in its own operational context. A vendor’s conformity paperwork does not discharge them, because the vendor never sees how the system is configured in the deployer’s hiring funnel or loan queue. Penalties for breaching operator obligations can reach €15 million or 3% of worldwide annual turnover, and harm to an individual may open separate civil claims under national law.
Article 25 marks the point where a deployer inherits provider duties as well: putting its own name on a high-risk system, substantially modifying it, or changing its intended purpose. Legal teams who read the Act only from the provider’s side tend to miss that the deployer’s own configuration choices can trigger it.
The consequence is procedural. Legal exposure created by a system the organisation cannot inspect can only be managed with evidence the vendor supplies, which makes algorithmic accountability a contractual and operational matter long before it becomes a legal one. The NIST AI Risk Management Framework reaches a similar conclusion from the risk side: its Core organises AI risk work into four functions, govern, map, measure and manage, and treats governance as a cross-cutting function that runs through the other three across the whole system lifecycle NIST AI Risk Management Framework (NIST AI RMF Core).
Vendor AI Governance and Due Diligence: Four Risks Standard TPRM Cannot See
Standard third-party risk management (TPRM) tests whether a service is available, secure and delivered as contracted, and none of those tests can see the four risks that AI systems add.
Algorithmic harm is the first. A vendor model can produce discriminatory outputs while every security control passes, because bias lives in the relationship between training data and the population the deployer applies the model to, and no penetration test touches that relationship. The organisation that deploys the model inherits the outcome for every person the model scores.
Model drift is the second. A vendor retrains, swaps a component or adjusts a system prompt between two contract reviews, and the behaviour the buyer validated at signature no longer describes the product in production. Nothing in the vendor’s uptime report changes, and the buyer’s file still shows the last approved state.
Supply chain opacity is the third. Many AI vendors build on foundation models from upstream providers, and the training corpus behind those models may be undescribed even to the vendor. The OECD’s due diligence guidance for responsible AI notes that risks can arise at any point along the AI value chain, from the people who label and moderate datasets to the deployers of finished systems (OECD.AI), and the AI supply chain is the reason the buyer often cannot trace them.
Emergent capability risk is the fourth. A vendor adds an AI feature, such as summarisation, scoring or automated suggestions, to a SaaS product the organisation already owns, and no procurement event marks the change. The contract, the risk tier and the approval all still describe a product without the feature.
Each of the four has an answer somewhere in a working programme:
| Risk category | Where the answer sits |
|---|---|
| Algorithmic harm | Key Components of an AI Vendor Due Diligence Framework |
| Model drift | Monitoring and Auditing AI Vendor Compliance Over Time |
| Supply chain opacity | Scrutinizing Vendor Subprocessor Lists and Data Handling Practices |
| Emergent capability risk | Integrating AI Vendor Governance into Third-Party Risk Management Programs |
Key Components of an AI Vendor Due Diligence Framework
A complete AI vendor due diligence framework asks for evidence in six categories that security questionnaires never reach: a substantive model card, training data provenance, bias testing methodology, AI-specific incident terms, change notification duties, and infrastructure certification kept in its own lane as evidence of infrastructure security alone.
ISO/IEC 27001 and SOC 2 certification, including the SOC 2 Type II report that covers a period of operation, answer whether the vendor protects information and runs its controls consistently. A reviewer who accepts them as the whole file has documented the wrong risk. Neither standard asks how the model behaves for a given population, what data taught it, or whether the vendor will say when either changes. ISO/IEC 42001, the AI management system standard, does address supplier management for AI, and a vendor holding it has at least described its own governance; a certificate still says nothing about a specific model’s behaviour in the buyer’s use case.
The NIST AI RMF Playbook supplies a practical vocabulary for the requests. It offers suggested actions aligned to each subcategory of the Govern, Map, Measure and Manage functions, and it is explicit that organisations may borrow as many or as few suggestions as fit their use case Measure and Manage (NIST AI RMF Playbook). An AI Bill of Materials, an inventory of a system’s model components, datasets and upstream dependencies, is a convenient container for the provenance evidence a vendor can already produce. An AI Governance Assessment of your own vendor programme shows which of the six categories your current questionnaires never ask about.
Model Card Requirements: Substance vs. Presence
A model card earns a place in the vendor file only when it carries the model’s purpose, performance figures, evaluation datasets, known limitations and bias results; a card that describes general capabilities is marketing with a template.
The test is to read the card against a use case the organisation actually runs. If the card reports accuracy for a population that bears no resemblance to the buyer’s applicants, patients or customers, it documents a different deployment. Adequate evidence names the evaluation datasets and their composition, gives performance figures per relevant subgroup where the use case affects people, states the conditions under which the model is known to fail, and carries a date and a version so the reviewer knows which model the card describes.
Red flags are easy to spot once the reviewer knows to look for them. Benchmarks with no dataset named, limitations written as generic disclaimers, a card that predates the model version in production, and a card that describes the vendor’s foundation model but says nothing about the fine-tuned version the buyer receives. A vendor that produces a card only after a request often has no internal document to draw from, so the file also records how quickly and from whom the card arrived.
Reviewers who lack ML expertise can still apply the test, because it depends on comparison rather than evaluation. The question is whether the card lets the buyer predict how the model will behave on the buyer’s own cases. If it does not, the presence of the document changes nothing about the risk.
Training Data Provenance and Legal Rights
Training data provenance evidence tells the buyer where a model’s data came from, on what legal basis the vendor holds it, and whether any of it could taint the outputs the buyer relies on.
Ask for the data source categories, the licensing basis for each, any use of web-scraped material, and the treatment of personal data in the corpus. Web-scraped content raises copyright questions and, where it contains personal data, questions about a lawful basis under GDPR. A vendor that answers with the phrase “publicly available data” has described accessibility, and lawful use of accessible data is a separate matter.
Adequate evidence is specific enough to be checked: named dataset families or licence categories, a description of how opt-out signals and takedown requests are handled, a statement of whether customer data from other deployments enters training, and the retention period for the corpus. An AI Bill of Materials that lists datasets alongside model components serves as the record. Where the vendor builds on an upstream foundation model, the adequate answer names the upstream provider and states which provenance claims the vendor can make itself and which it inherits.
The red flag is a vendor unable to describe its own corpus. That is common when the vendor consumes a foundation model through an API, and it is informative: the provenance gap is real, and the file should record it as an accepted or unaccepted risk rather than leave the field blank. The exposure lands on the deployer when an output reproduces protected material or when a regulator asks how a decision-influencing model was trained.
Bias Testing Evidence Standards
Bias testing evidence is adequate when the vendor documents its methodology, the fairness metrics it measured, the breadth of its evaluation datasets and how often it retests, since a statement that testing occurred proves nothing about what was tested.
A request for bias testing methodology should return four things: the metrics, the groups and intersections evaluated, the evaluation data with its provenance, and the date of the last run and the cadence for the next. Vendors who test once at launch and never again leave the buyer holding evidence about a model that no longer exists, so the periodic refresh obligation belongs in the evidence standard as well as in the contract.
Fairness Metrics Named in the Evidence
Two metrics recur in credible evidence. Demographic parity compares the rate of favourable outcomes across groups, and equalized odds compares true positive and false positive rates across groups. The two generally cannot both hold when underlying base rates differ between groups, so a vendor that reports both as satisfied is either working with unusual data or has not looked closely.
The useful question for the buyer is which metric the vendor chose and why. A hiring tool judged on demographic parity answers a different fairness question than one judged on equalized odds, and the buyer’s legal team may need one of them for its own jurisdiction. Evidence that names the metric lets the deployer decide whether the vendor’s choice matches its own obligation.
Evaluation Set Coverage and Refresh
Evaluation coverage decides whether the metrics mean anything. A test set drawn from one country, one language or one demographic mix says little about a deployment in another. Adequate evidence describes who is represented, how intersections such as age and gender together were handled, and whether the set reflects the population the buyer will actually score.
Refresh matters for the same reason drift matters: the model, the data and the population all move. Vendors who commit to retesting after each material model update and at a stated interval give the buyer an evidence trail. Vendors who offer a single historical report give the buyer a snapshot with no expiry date.
AI-Specific SLA and Change Notification Terms
An AI vendor SLA covers model behaviour failures with a 24 to 72 hour notification window for unexpected outputs and bias incidents, and it pairs that incident clock with a contractual notice standard for changes to the model, the features and the vendor’s own suppliers.
Uptime guarantees measure availability, and a biased model is fully available. The AI-specific terms therefore define what counts as an incident: outputs outside the documented performance envelope, a bias finding the vendor’s own monitoring surfaces, an unexpected change in output distribution, and a security event involving training or inference data. Notification within 24 to 72 hours gives the buyer time to pause automated decisions before the effect compounds.
The change notification standard is a different instrument. It obliges the vendor to notify the buyer at least 30 days before three kinds of event: model updates, new AI feature releases and subprocessor changes. Thirty days is the shortest window in which a buyer can re-test, re-classify and, where needed, brief affected staff. Anything described as “reasonable notice” leaves that window to the vendor’s discretion.
Adequate evidence is the clause itself, tied to defined events, together with a change log the vendor maintains and shares. Red flags include notice only of “material” changes with no definition, notice after deployment, notice limited to security patches, and subprocessor notification limited to the parties named in the data processing agreement. Later parts of a working programme decide what happens when notice is late or missing; this standard fixes only the period and the three events it reaches.
Scrutinizing Vendor Subprocessor Lists and Data Handling Practices
Scrutinizing an AI vendor’s subprocessor list means tracing the full chain from the SaaS product the organisation buys to the foundation model that generates its outputs, since every upstream party shapes results without appearing in the buyer’s contract.
Buyers who diversify across AI vendors often discover they have diversified very little. Fourth-party AI risk is the exposure created by the parties behind the vendor: the foundation model provider, the fine-tuning provider, the annotation service. It differs from classic fourth-party risk because these parties determine what the product does, and not only whether it stays online.
Mapping the AI Supply Chain from SaaS to Foundation Model
The AI supply chain runs from the enterprise to an AI SaaS vendor, from the vendor to a foundation model API from OpenAI, Anthropic, AWS Bedrock or Google Vertex, and from that model to a training corpus of web-scraped, licensed and sometimes customer-derived data.
Each layer is a governance gap. The first request is a complete subprocessor list. That means more than the sub-processors named in the data processing agreement (DPA): it includes every upstream AI provider, including foundation model vendors, fine-tuning providers and annotation services. Partnership on AI’s five-step path for responsible AI supply chains starts with adopting data enrichment guidelines and introducing internal governance, and it notes that the treatment of data enrichment workers shapes the quality, safety and reliability of the resulting systems (Partnership on AI). Annotation services belong on the list for that reason.
A hosting layer is a different thing from a model provider. AWS Bedrock and Azure OpenAI Service are cloud services through which models are accessed; a vendor that answers “we run on Azure” has told the buyer where the compute sits and nothing about whose model generates the output. Ask for both answers separately.
Most enterprise AI SaaS is built on the same three or four foundation model providers, an estimate in the range of 70 to 80 percent. Spreading spend across many AI SaaS vendors therefore migrates the concentration upstream, and AI supply chain concentration risk shows up only when the portfolio is drawn against its dependencies. A failure, a policy change or a pricing change at one provider then reaches every vendor built on it at once.
The provider swap is the most overlooked event in the chain. When a vendor moves from one model developer to another, or replaces the model behind an unchanged API endpoint, the DPA’s named parties may stay identical while every output property changes. A list of DPA sub-processors never records that event, so the notice standard has to name a foundation model swap explicitly as a change it reaches.
Concentration Matrix for the Vendor Portfolio
A concentration matrix draws each AI vendor against the foundation model provider behind it, the hosting layer it uses and the fine-tuning or annotation parties it relies on, and it flags any provider that carries more than one vendor.
Build it from the subprocessor disclosures, and leave a cell empty when the vendor will not answer, since an empty cell is itself a finding. An example with generic vendor types and placeholder provider labels shows the shape:
| AI vendor in portfolio | Foundation model provider | Hosting layer | Fine-tuning / annotation party | Concentration flag |
|---|---|---|---|---|
| Candidate screening tool | Provider A | Cloud host 1 | Annotation firm X | Shared with two others |
| Support assistant | Provider A | Cloud host 2 | Undisclosed | Shared with two others |
| Contract analytics | Provider B | Cloud host 1 | Vendor-internal | Single |
| Sales forecasting | Provider A | Cloud host 1 | Undisclosed | Shared with two others |
Three of four vendors resting on Provider A means that an outage, model withdrawal or policy change there interrupts three business processes together. Ranking by the size of the flag rather than by contract value tells the review team which upstream relationship deserves an exit plan, and which vendors’ disclosure gaps hurt the most.
Data Isolation Verification and Training Data Use
Whether a vendor can use the buyer’s data to train its model depends on the contract terms, the technical isolation of customer data, and the retention and logging of inference data, all three of which the buyer should verify rather than accept from a summary page.
Four questions structure the verification. Does the contract prohibit use of customer data for training, including by upstream providers? Is customer data technically isolated from other tenants, which is what GDPR data isolation depends on in practice? How long does the vendor and each upstream party retain prompts and outputs? And who holds the inference logs and can read them? A vendor that answers the first question yes and cannot answer the fourth has given the buyer half a control.
Isolation and monitoring pull in opposite directions, and vendors’ own offers show the conflict. Anthropic’s Enterprise Frontier Safeguards, announced on 1 September 2026, pairs zero data retention with misuse detection by storing data in cloud infrastructure controlled by the customer, on the reasoning that sophisticated misuse spans many sessions and accounts and that detection needs data kept for a meaningful period Enterprise Frontier Safeguards (Anthropic). The design choice generalises: a buyer who demands zero retention gives up vendor-side detection, and a buyer who accepts retention needs to know whose infrastructure holds the data.
Ask the vendor to show the retention schedule, name the party that controls the logs, and describe how model training use is excluded technically and not only contractually. Where the contract allows, a small experiment gives more confidence than any statement: place a uniquely identifiable test record in a non-production tenant and ask the vendor to show where it persists and for how long. The result tells the reviewer whether model training data use is prevented by architecture or by promise.
What Should a Subprocessor Change Notice Contain?
A subprocessor change notice should name the new party, the function it performs, the data it will touch and the date the change takes effect, so the buyer can update the concentration matrix without chasing the vendor for detail.
A notice that says only “we have added a sub-processor” leaves the reviewer to guess the rest. Ask for the legal name of the party, its role in the chain (hosting, foundation model, fine-tuning or annotation), the country where it processes data, and whether it will receive prompts, outputs or training material. Ask also whether the change alters which model generates the output, because that is the case a DPA list never records.
The notice should arrive early enough for the buyer to object. A buyer that learns of the change after it takes effect can only record it. The contract should therefore say what the buyer may do on objection: pause the affected use, require the vendor to keep the old party in place, or leave the contract under the exit terms.
Integrating AI Vendor Governance into Third-Party Risk Management Programs
Integrating AI vendor governance into a third-party risk management programme means keeping the existing intake, assessment, contracting and monitoring workflow and grafting AI modules onto it at four positions, with AI system risk classification promoted to a primary tiering variable.
Tiering by spend is where the graft usually fails. A $5,000-a-year AI-powered hiring tool can carry more EU AI Act exposure than a $500,000 infrastructure contract with no AI in it at all, and a programme that tiers on spend, operational criticality and data sensitivity puts the second vendor under closer review than the first. The fix is to add a fourth variable, and to let it escalate a vendor irrespective of the other three.
Nothing new is built. PwC’s responsible AI and TPRM guidance takes the same position, extending the existing programme with AI-specific considerations. The AI modules attach to the vendor onboarding workflow at these handoff points:
- Pre-contract disclosure. A short questionnaire asks whether the vendor uses or embeds AI, and which functions. A yes triggers classification.
- Assessment substitution. For any vendor classed as AI, the AI-extended questionnaire replaces the standard security one, drawing on the evidence set the due diligence framework defines.
- Contract clause insertion. The procurement owner attaches the AI schedule to the draft paper at the step before legal review, so legal sees the schedule as part of the first draft.
- Monitoring calendar entry. The vendor enters the monitoring calendar with its event triggers as well as its review dates.
- Reclassification at renewal. The tier is recalculated from current inputs and the result is logged.
An AI Governance Assessment shows which of these five handoff points your current workflow lacks, and it usually finds the disclosure gate missing first.
AI Risk Classification Overlay for Vendor Tiers
The AI risk classification overlay adds a fourth tiering variable, the vendor’s AI system risk classification, and sets the vendor’s tier at the highest level any of the four variables produces.
Traditional variables stay as they are: spend, operational criticality, data sensitivity. The overlay classifies the AI function by the highest-risk use it supports, using the EU AI Act’s high-risk classification where the use falls in a regulated area such as employment, credit or access to essential services, and the NIST AI RMF’s view of impact on people where it does not. A vendor whose AI function falls in a high-risk category goes to the top tier regardless of spend, because the exposure attaches to the use and the contract value does not change it.
Alignment with ISO/IEC 42001 supplier management helps here. The standard expects an organisation to control the AI-related products and services it acquires, and the overlay gives that expectation a concrete owner: the tier assignment. Teams already certified against it can map their supplier controls to the overlay instead of writing a second procedure.
The overlay sorts the vendor portfolio differently from the old key. The hiring tool jumps from the lowest tier to the highest, the infrastructure contract stays where it is, and reviewers who had spent their hours on the largest contracts find themselves reading model cards for the smallest ones. That reordering is the intended effect, and it should be expected to draw objections from the contract owners of the vendors that moved.
Trigger Events and Escalation in AI Vendor Monitoring
A trigger event moves a vendor’s tier: the signal reaches the vendor’s TPRM analyst and business owner, the vendor moves to the next higher tier, and the head of third-party risk signs the new tier before any other work begins.
The order matters. If the signal reaches only the business owner, the tier stays where it was and the vendor continues on its old review calendar. If the analyst moves the tier without a signature, the change cannot be defended in an audit. Naming the recipient, the destination and the signer turns a vague escalation into a workflow step.
At renewal the same mechanism runs on a schedule. Reclassification is a tier recalculation: the inputs are the current spend, criticality and data sensitivity, the vendor’s present AI function list as disclosed at renewal, and the classification of each function. The output is a tier, and the tier change or confirmation is entered in the vendor’s log with the date, the inputs used and the signer. A vendor that added an AI feature during the contract term will therefore appear at renewal in a tier its original paper never anticipated.
The catalogue of events that count as triggers, and the response each one draws, is defined by the monitoring programme. What the integration needs from it is a single fact per event: which tier the vendor moves to, and who signs.
What Belongs in the AI Contract Schedule?
The AI contract schedule is one attachment to the vendor paper that gathers every AI-specific term in one place: the notice standard, the incident clock, the audit right, the remediation terms and the exit rights.
Keeping these terms in a schedule means the procurement owner attaches the same document to each AI vendor’s draft, and legal reviews one known text instead of hunting for clauses spread across a master agreement. Vendors that push back can be shown which terms are fixed by the tier and which are open to negotiation.
The schedule also gives the monitoring calendar something to read. The procurement owner attaches the schedule at the step before legal review, and the vendor’s analyst reads it again when the vendor enters the monitoring calendar. Each clause names an event or a period, such as the change notice or the incident window, and each one becomes an entry the analyst can track. When the tier changes at renewal, the schedule attached to the vendor moves with it, so the terms always match the tier the vendor currently holds.
AI Vendor Governance Roles: Who Owns Third-Party AI Accountability
Ownership of vendor AI governance belongs to no single function, so the workable answer assigns the review across procurement, legal, IT and security, data governance and an AI ethics lead, with legal and the responsible AI lead each holding a veto over approval.
Every function in that list is partly right about why the work belongs to someone else. Procurement says the contract is the control point. Legal says liability is a legal matter. Security says it already reviews vendors. The decision that has to come out of the argument is who signs, and the answer is a shared set of signatures with named vetoes.
Why No Single Office Can Own the Whole Review
Each of the five functions holds a genuine and non-transferable piece of the evidence, and the piece each one cannot supply is exactly the piece another one already has.
Procurement controls the contract paper and the commercial leverage, and cannot judge whether a model card is substantive. Legal reads the EU AI Act and the deployer’s exposure, and cannot say whether a bias metric was appropriately chosen. IT and security validate integration security and model documentation, and rarely hold the licensing history of a training corpus. The data governance function knows lineage, retention and lawful basis, and does not evaluate fairness. The AI ethics function reads bias and fairness evidence, and has no contract to negotiate with.
Assigning the whole review to one office therefore leaves one of the five pieces to a reader who cannot interpret it. The NIST AI RMF says as much about the work itself: its functions are meant to be carried out in ways that reflect diverse and multidisciplinary perspectives, including views from outside the organisation, because a diverse team surfaces problems and emergent risks a single perspective would miss NIST AI Risk Management Framework (NIST AI RMF Core).
The practical failure is familiar. Procurement sends the questionnaire, security returns a certificate summary, legal reviews the indemnity, and nobody reads the bias evidence because it belongs to nobody’s checklist. The gap between functions is where the file goes quiet.
RACI for Approving a New AI Vendor
The RACI matrix for a new AI vendor gives procurement the disclosure questionnaire and the contract AI terms, gives legal the regulatory review and sign-off before execution, and puts technical validation, data provenance and bias review with the functions that hold the expertise.
| Activity | Procurement | Legal | IT / Security | Data governance | AI ethics |
|---|---|---|---|---|---|
| Disclosure questionnaire and contract AI terms | R | A | C | I | C |
| SLA terms, EU AI Act obligations, deployer liability review, sign-off before execution | C | R, A | C | I | C |
| Technical validation: model documentation and integration security | I | I | R, A | C | C |
| Training data provenance and data isolation verification | C | C | C | R, A | I |
| Bias and fairness evidence review | I | C | C | C | R, A |
R marks who does the work, A who answers for it, C who is consulted, and I who is informed. Procurement is responsible for initiating the questionnaire and the AI terms, and accountable to legal for regulatory alignment.
Two roles carry vetoes. Legal and the AI ethics function each hold an independent veto over an AI vendor approval. Procurement cannot override either one without a documented exception process, in which the exception, the reason and the approving executive are recorded in the vendor approval workflow. Without the process, commercial pressure turns the veto into a formality.
Where no dedicated responsible AI function appears on the organisation chart, bias and fairness review defaults to the CISO or the Chief Data Officer, whichever office the organisation trusts with evidence quality. It lands there by default, and the organisation should say so explicitly.
CISO or Chief AI Officer: Who Inherits the Sign-Off
In an enterprise without a Chief AI Officer, the Chief Information Security Officer (CISO) inherits third-party AI accountability by default, along with a capability gap that nobody assigned deliberately.
The debate between the two roles is a debate about capability. A CISO’s training covers threats, controls and incident response, and bias evaluation and training data rights fall outside it. A Chief AI Officer would own model behaviour, but where the role does not exist, the security function absorbs the work because vendor risk already sits with it. The result is sign-offs by an office that can validate the infrastructure and cannot evaluate the model.
An interim AI Vendor Review Board is the structure that holds the decision while permanent accountability is settled. Its members are the five functions in the matrix, with the CISO or Chief Data Officer covering the responsible AI seat when it is vacant. It meets when a new AI vendor or a reclassified one needs approval, it holds the two vetoes, and it keeps the exception log. It also names the conditions under which it dissolves: the appointment of an accountable executive who then owns the sign-off.
The board is deliberately temporary. A standing committee with no owner tends to become the owner, and the capability gap it was created to cover then becomes permanent.
Vendor AI Governance vs Traditional Vendor Management: Key Differences
AI vendor governance differs from traditional vendor management along six dimensions, contract lifecycle, risk type, assessment evidence, liability, behaviour and monitoring, and each difference is categorical, because the vendor’s product can change without an amendment, an invoice or an outage.
Procurement will accept “more thorough” as a reason for a bigger checklist, and will reject it as a reason for a different process. The argument that wins the point runs through structure. In each of the six dimensions, the traditional process contains a step whose input no longer exists for an AI vendor.
Lifecycle, Risk Type and Assessment Evidence
A traditional vendor contract is a static statement of work reviewed annually, whereas an AI vendor relationship changes whenever the vendor updates a model, adds a feature or changes a subprocessor, so the contract lifecycle becomes living governance.
The risk question changes with it. A traditional review asks whether the service is available and secure. An AI review adds whether the model is performing as expected, whether it is drifting, and whether it produces biased outputs. For procurement, this is a difference in process design. A traditional purchase runs its acceptance test once, against a fixed build. An AI purchase needs an acceptance test that the buyer can repeat on demand, against outputs and not against a build.
The evidence changes with the risk. The traditional file holds a security questionnaire and certifications such as ISO/IEC 27001 and SOC 2. The AI file adds a model card, a bias report, a training data summary, incident history and a change notification log.
An algorithmic audit is a different exercise from a security audit. A security audit checks controls against a standard using assessors who read configurations, logs and policies. An algorithmic audit measures model behaviour on data, using statistical evaluation and needing people who can design test sets and interpret fairness metrics. The evidence types differ and so does the expertise, so the same auditor cannot be assumed to cover both.
Liability, Behaviour and Monitoring Cadence
Liability, behaviour and monitoring each remove an assumption on which the traditional review was built, and each can be carried into the comparison as a single contrast.
Liability contrasts two role vocabularies. The data processing agreement templates procurement already uses have two seats, controller and processor, and no seat for a deployer. A template built on that pair cannot record the difference between a controller and a deployer, so the AI terms need their own schedule that the template points to.
Behaviour contrasts a fixed deliverable with an undocumented one. A traditional supplier’s deliverable is fixed by the contract and moves only by amendment. Model conduct, including emergent behavior no one anticipated, is written down nowhere that either party holds. The amendment event on which the whole traditional review trigger was built has gone.
Monitoring contrasts dates with events. The traditional review runs annually or after an incident, while AI vendor monitoring is driven by events rather than by dates.
Read together, the three contrasts explain why procurement’s instincts fail in a predictable way. Each one assumes that a change to the supplier’s deliverable is visible to the buyer through a process the buyer controls. For AI vendors, none of the three assumptions holds by default.
Side-by-Side Comparison of Traditional and AI Vendor Management
The comparison table sets both approaches against the six dimensions so that a reviewer can see, in one place, which process step each difference removes or replaces.
| Dimension | Traditional vendor management | AI vendor governance |
|---|---|---|
| Contract lifecycle | Static statement of work, annual review | Living governance: model updates, new AI features and subprocessor changes keep the contract relevant |
| Risk type | Delivery and security: is the service available and secure? | Adds algorithmic risk: is the model performing, drifting or producing biased outputs? |
| Assessment evidence | Security questionnaire, ISO/IEC 27001, SOC 2 | Model card, bias report, training data summary, incident history, change notification log |
| Liability | GDPR data processor | Deployer, and in some cases provider, under the EU AI Act |
| Behaviour | Features defined in the contract, changed by amendment | Emergent behavior not documented anywhere; vendor can update the model between contracts |
| Monitoring | Annual review or incident-triggered | Event-driven |
The certification row shows a structural change in the assessment step. In the traditional process, that step takes one input, the certificate, and the certificate is enough to close it. In the AI process, that step takes a second input, model evidence, and the step has no closing rule until the buyer defines one. The step stays in the workflow, and its exit condition must be rewritten.
A useful way to run the comparison in a meeting with procurement is to work through a live vendor. Fill the table for one AI vendor already under contract, mark the cells where the current file holds no evidence, and count them. The result is a gap list built from the organisation’s own paper, which persuades better than any generic argument about difference.
Monitoring and Auditing AI Vendor Compliance Over Time
Monitoring an AI vendor runs on three signal streams, vendor notices, regulatory movement and public incident intelligence, and each stream needs three fixed properties: a named owner, a log location and a re-assessment rule that says which signal reopens the vendor’s file. These streams replace the review calendar the comparison section describes.
The vendor’s silence is the hardest signal to read, and the contract is the only place it can be turned into a finding. A stream with no owner produces no log entry, and a log entry with no re-assessment rule produces no action. The sections below set those three properties for each stream in turn.
Three Non-Calendar Escalation Triggers for AI Vendor Monitoring
Three signal streams should escalate monitoring beyond scheduled reviews: the vendor’s own contractual notices, movement in the regulatory landscape, and public intelligence about incidents involving the vendor.
The first stream is vendor-triggered events: notices of model updates, new AI feature releases, subprocessor changes and security or privacy incidents, all delivered under the notice standard the due diligence framework sets. Each one triggers something on the buyer’s side. A model update notice sends the organisation back to the model card, and to its own acceptance tests, to see what changed. A new feature notice moves the feature into the assessment queue. A subprocessor change updates the concentration matrix. An incident notice opens the buyer’s own incident record. If the vendor fails to provide notice, that failure is itself a material compliance failure, because it shows the vendor’s process cannot be relied on to surface the next change.
The second stream is regulatory. Expansions of the EU AI Act high-risk list, updates to NIST AI RMF guidance, and sector rules such as DORA operational resilience requirements for financial services or FDA AI guidance for healthcare can each turn a vendor’s AI function into one that needs elevated assessment. The NIST AI RMF Playbook is described as a living resource with updates about twice a year Measure and Manage (NIST AI RMF Playbook), which makes a standing calendar entry to check it a cheap trigger. When the landscape changes, re-evaluate whether the vendor’s systems now need a higher level of assessment.
The third stream is market intelligence: regulatory actions, published bias research, customer lawsuits. These are often the earliest signals of an AI system’s problems, and they typically surface before the problem reaches the buyer’s own deployment. The OECD tracks and categorises risks through its AI Incidents and Hazards Monitor (OECD.AI), which gives the monitoring team a public feed to search by vendor and function.
Continuous compliance monitoring needs an owner for each stream, a place where signals are logged, and a rule for when a signal starts a re-assessment.
Negotiating and Exercising AI Vendor Audit Rights
AI vendor audit rights should include the right to request an independent third-party audit of AI system performance and bias characteristics every 24 months for high-risk systems, and the buyer has to be able to exercise that right in practice.
A right that cannot be used leaves the buyer with paper. The negotiation therefore settles who qualifies as an independent auditor, what data the auditor may see, how confidential material is protected, and who pays. Vendors commonly resist on confidentiality grounds, and the objection deserves a technical answer.
Audit Scope, Frequency and Auditor Independence
The audit scope covers model performance on the buyer’s use case and bias characteristics on the metrics the vendor has already documented, so that the audit tests the vendor’s own claims. A 24-month interval applies to high-risk systems, and a trigger event can bring an audit forward.
Independence means the auditor has no commercial relationship with the vendor that would compromise the finding, and the contract should say how such a person is chosen. Without that clause the vendor nominates its own reviewer, and the audit becomes a second self-assessment.
Verifiable Audit Without Model Disclosure
Fairness as a Service, a published protocol for verifiable and privacy-preserving fairness auditing of machine learning systems, shows that an audit need not expose the model or the data. It is model-agnostic, independent of any specific fairness metric, uses zero-knowledge proofs to assure that the audit steps were performed correctly, and in experiments scaled to large auditing scenarios of more than 1,000 participants (Toreini, Mehrnezhad and van Moorsel).
The relevance to negotiation is practical. A vendor that says a bias audit would reveal its intellectual property has an answer available from published research, and the buyer can propose an audit protocol that preserves confidentiality. The objection then rests on willingness and no longer on feasibility.
Telemetry, Observability, and Bias Monitoring Requirements
The buyer should require inference-level telemetry, or outcome data sufficient to support independent bias monitoring, so that the organisation can measure the vendor’s model on its own decisions and does not rely on the vendor’s summary.
Inference telemetry means records of inputs, outputs and confidence for the buyer’s own use of the system, retained for long enough to compute fairness measures over time. Outcome data means what happened after the decision, such as which candidates were hired or which applicants defaulted. Bias monitoring compares selected metrics across groups on this data, and it can detect a change in behaviour without needing to know the cause.
The EU AI Act supports the arrangement from the provider side. Article 72 requires providers of high-risk systems to establish a post-market monitoring system that actively and systematically collects, documents and analyses relevant data on performance throughout the system’s lifetime, and that data may come from deployers (EU AI Act Article 72). The Explorer lists the article as coming into force on 2 August 2026. A deployer’s telemetry is therefore an input the vendor needs anyway, and the contract should set out how the data flows in both directions.
The observability requirement should also name the format, the retention period and the access route, because telemetry that sits in a vendor’s logging system without a defined export gives the buyer nothing when an incident occurs.
Remediation Clocks, Exit Rights and Undisclosed Feature Detection
Remediation terms give the vendor 30 days to correct a bias finding above the defined threshold, require a written vendor remediation plan, and give the buyer termination rights for material undisclosed model changes.
The 30-day clock starts when the finding is confirmed, and the plan carries dates for each corrective action. The buyer’s evidence that the fix worked is a repeat measurement on the same metric, since a vendor’s statement that it addressed the finding tests nothing. If the vendor misses the date, escalation follows the terms already agreed.
Termination rights for material undisclosed model changes address the case where the buyer discovers a change nobody notified. The right means the buyer can leave without penalty, and the exit terms should cover data return and deletion, transition support and the buyer’s ability to keep operating on an interim basis while replacing the system.
Undisclosed AI features in existing SaaS products need their own detection routine. Two mechanisms work together: a periodic review of the vendor’s product release notes, which often announce AI features in marketing language before any contract change, and a contractual AI feature disclosure requirement at each renewal cycle. Reviewing release notes turns the vendor’s own publicity into a monitoring input at almost no cost.
Common Challenges in AI Vendor Due Diligence and How to Overcome Them
AI vendor due diligence most often fails through five recurring gaps, documentation theater, certification substitution, scope blindness, legal ambiguity and cadence failure, and each one shows a symptom in the vendor file that a reviewer can learn to recognise.
A shared root cause sits under all five: the organisation holds little contractual leverage, so minimal disclosure stays the vendor’s cheapest option for as long as silence costs nothing. The Future of Life Institute’s Winter 2025 AI Safety Index shows how uneven disclosure is even among frontier developers: of eight companies graded, the best, Anthropic, earned a C+ overall with a score of 2.67, and five of the eight fell to D or below AI Safety Index (Future of Life Institute). Vendors downstream of those developers may have less to disclose than the developers themselves.
Documentation Theater and Certification Substitution
Documentation theater and certification substitution both give the reviewer a complete-looking file that holds no evidence about the model.
- Documentation theater.
- Symptom: the vendor provides a model card that describes general capabilities but contains no performance metrics, evaluation datasets, bias testing results or known limitations.
- Root cause: nothing obliges the vendor to provide substantive documentation, so it provides the cheapest document that closes the request.
- Response: set specific contractual disclosure fields: performance benchmarks on standardised evaluation sets, bias testing methodology and results, and a training data summary. A vendor that fills them has produced evidence; a vendor that refuses has produced a decision input.
- Certification substitution.
- Symptom: the reviewer marks the AI vendor’s file complete after the first track of evidence arrives, and no second track exists to hold the remaining work.
- Root cause: risk teams lack AI-specific evaluation criteria, so they apply the criteria they have, and their workflow tool offers one checklist per vendor.
- Response: run a two-track assessment. The workflow opens an infrastructure track and a model-behaviour track for every AI vendor. Each track keeps its own list of open items, the tracks run side by side, and neither track can close the other’s finding.
The two-track design also settles an internal argument. A reviewer who wants to close the infrastructure track can do so and still cannot close the file, because the model behaviour track shows open items. The file closes only when both lists are empty.
Scope Blindness, Legal Ambiguity and Cadence Failure
Scope blindness, legal ambiguity and cadence failure hide in the gaps between programme documents: the inventory, the contract and the review calendar.
- Scope blindness.
- Symptom: the vendor inventory lists “CRM” or “HR platform” and flags nothing about embedded AI features such as predictive scoring or automated decision suggestions.
- Root cause: inventory entries are labelled by product category, and a label such as “CRM” records what the vendor was bought for, not what the product now does. The label hides the AI function.
- Response: run an inventory-relabelling exercise. Export the inventory. For each entry, take the function list the vendor has supplied and mark every function that scores, ranks, predicts, generates or suggests. Set that list beside the product-category label. Where the list shows a function the label does not name, correct the label so it states the AI function, and route the entry into the assessment queue. The next pass then starts from the corrected label. Detection of new AI features stays with the monitoring programme.
- Legal ambiguity.
- Symptom: contracts stay silent on whether fine-tuning, substantial customisation or rebranding flips the enterprise’s own regulatory status.
- Root cause: nobody has checked the existing contracts for customisations, so nobody knows which ones crossed the deployer-provider boundary.
- Response: run a contract-inventory audit. List every AI contract in which the organisation fine-tuned, configured, rebranded or repurposed the vendor’s system. Give each entry to legal for a decision on whether the customisation crossed the boundary, and record the decision in the contract file. Where an entry crossed, open a remediation item with an owner. Legal review then comes before any new AI customisation or rebranding proceeds.
- Cadence failure.
- Symptom: the vendor was last assessed 18 months ago; since then it has updated its foundation model twice and shipped three new AI features.
- Root cause: nothing links the vendor’s release history to the assessment record, so the old assessment still shows as current.
- Response: compare the date of each assessment with the date of the vendor’s last model update, and route every assessment that is older into the assessment queue. Event-triggered reassessment, which the monitoring programme describes in full, keeps the queue current afterwards.
Because these three failures live between documents, a single owner rarely sees them. A short AI governance gap analysis that walks one vendor through inventory, contract and calendar in sequence tends to reveal all three within an afternoon, and gives the review lead something concrete to bring to the next steering meeting.
Measuring Vendor AI Governance Effectiveness with KPIs and Scorecards
A vendor AI governance scorecard measures effectiveness in three tiers, coverage, quality and outcomes, because completion counts alone show that the process ran and never show whether it found or fixed any risk.
A 100 percent completion rate on a questionnaire that never asked about model behaviour scores perfectly and protects nothing. Most programmes count activities such as assessments completed and contracts reviewed, and report those counts as effectiveness. Only the outcomes tier can answer the question a board eventually asks: whether the programme works.
Tier 1: Coverage Metrics and Shadow AI Detection
Coverage metrics show how much of the AI vendor estate the programme sees: the percentage of third-party AI tools with completed due diligence, the percentage on the inventory, and the percentage on a defined review cycle.
The target is 100 percent of identified AI vendors assessed. The word “identified” carries the weight, which is why the tier includes a shadow AI count as the inverse coverage indicator: the number of AI tools found in use that the programme did not know about. Shadow AI detection compares what the inventory says with what network, identity and spend data show, covering sign-in records to AI services, OAuth grants to third-party applications and expense lines for AI subscriptions. A periodic scan should surface fewer than 5 percent shadow AI tools, and a higher rate means the inventory is missing more than the scan can absorb.
Shadow AI in SaaS deserves attention because it enters through the ordinary channels of work. An employee enables an AI assistant in a tool the organisation already approved, or a team signs up for a service on a corporate card. Neither event passes through procurement.
The measurement trap sits in the coverage number itself. A programme can raise coverage by adding easy, low-risk vendors to the denominator, and coverage then improves while exposure does not. Report coverage of the high-risk tier separately from coverage overall, and treat the shadow AI count as the check on both.
Tier 2: Quality Metrics and the Completion Trap
Quality metrics test the substance of assessments and not merely their completion, and governance quality metrics should show whether the file for each vendor contains evidence about the model.
Three measures carry the tier. The first is the percentage of vendor assessments that contain substantive model documentation, as opposed to a security certification alone; the target is more than 80 percent. The second is the percentage of AI vendor contracts with change notification clauses. The third is the percentage of high-risk AI vendors with bias testing evidence reviewed, where the target is 100 percent with bias evidence on file. This is the bias testing evidence rate. A related measure for high-risk vendors is the EU AI Act compliance rate: the share whose deployer obligations have a documented owner and a documented status.
Substance needs a definition before it can be counted. A workable rule: the assessment counts as substantive only if a reviewer, using the model card and the bias evidence, could predict the model’s behaviour on one of the organisation’s own use cases. Assessments that fail the rule do not count, however complete the questionnaire.
The quality tier shows a programme’s real state when coverage looks strong. Coverage can reach 100 percent while the quality rate sits at 30 percent, which means the programme has looked at every vendor and learned little about most. That combination is compliance theater at scale.
Tier 3: Outcomes Metrics and the Feedback Loop
Outcomes metrics measure whether the programme changed what happened: the AI vendor incident rate, the share of incidents that were pre-identified as risks in the assessment, the mean time to detect undisclosed model changes, and the audit finding closure rate.
Quarterly reporting on all three tiers, with trend lines, shows whether quality and outcomes improve as coverage matures. A programme whose coverage climbs while outcomes stay flat is measuring effort.
Incident Attribution: Missed Risk or Unfixed Risk
When an incident lands, the diagnostic question is whether the assessment failed to detect the risk or detected it and the mitigation failed to fix it. The two verdicts call for opposite responses: an assessment that never saw the risk needs more coverage or better questions, and a risk that was seen and left open needs better follow-through.
The metric is the percentage of vendor incidents that were identified as risks in the assessment beforehand. A low figure points to assessment quality, and a high figure paired with recurring incidents points to remediation. A scorecard that cannot separate the two cannot tell a programme that needs more reach from one that needs more discipline.
Detection Time and Closure Rate
Mean time to detect undisclosed model changes measures how long the vendor’s behaviour differed from the approved state before the buyer noticed. It rewards the monitoring streams and the telemetry that make changes visible, and it penalises reliance on vendor notices alone.
The audit finding closure rate, with a target of closing findings within 90 days, measures follow-through on what audits and reviews turn up. The two measures together show whether the programme sees change quickly and acts on what it sees, and the trend in each shows whether the answer is improving.
How Do You Decide Whether to Approve, Condition or Reject an AI Vendor?
Approve an AI vendor when its evidence meets the adequacy bar for its tier, approve with time-boxed conditions when a gap can be closed by a dated commitment, and reject when the gap sits in something the vendor cannot or will not disclose.
The three outcomes turn on a small set of questions that the earlier work produces, and the review board can answer them in sequence.
| Outcome | When it applies | Who signs | What follows |
|---|---|---|---|
| Approve | Evidence meets the bar for the vendor’s tier and no unresolved veto | Board, with legal and AI ethics concurring | Vendor enters the monitoring calendar |
| Approve with conditions | Gap is closable, the vendor has committed in writing to a date, and the use case can run in a restricted mode meanwhile | Board, with the condition and its expiry recorded | The expiry date is a monitoring entry; a missed date reverts the decision |
| Reject | Gap sits in provenance, behaviour or disclosure the vendor cannot or will not supply for a high-risk use | Board, with the exception process closed | Business owner told which evidence would reopen the case |
Concentration changes the plan without changing the verdict. A vendor that shares an upstream provider with two others can still be approved, and the review adds an exit or fallback plan for that provider. Conditional approvals need discipline: a condition without an expiry becomes a permanent approval by inertia, so the expiry belongs in the same log as the decision.
A rejection is easier to defend when it names the missing evidence. “The vendor could not describe its training corpus for a hiring use” gives the business owner something to go back to the vendor with, and it leaves the door open if the vendor later supplies the evidence.
How Do You Start Without Rebuilding the Programme?
Start with one small experiment: pick the three vendors whose AI outputs touch people or money, run the extended assessment on them, and read what the results show about the existing programme before changing anything else.
The point of a small first move is the information it produces. A programme rebuilt on assumptions carries the assumptions forward, while three assessments produce observations the team can weigh.
A workable first cycle has four steps. First, scan the vendor inventory and the product release notes for embedded AI, and list the vendors that use it. Second, sort the list by what the AI decides: hiring, credit, pricing, service access, customer-facing text. Third, run the AI-extended assessment on the top three and file the results next to the standard security file. Fourth, compare the two files and note what the standard file never held.
Three observations decide what happens next: how many evidence categories the vendors could not answer, how long the extended assessment took compared with the standard one, and whether any vendor’s answers changed how the team classified its use. Each is an observable result, and none needs a budget request.
The next move belongs to the team that ran the experiment. If the gaps cluster in one evidence category, that category is the place to invest. If the assessments went smoothly, the team can widen the sample. If a vendor could not answer basic questions, the approve, condition or reject path applies to that vendor. Progress is measured by what the small cycle taught the team, and a wider rollout can follow when the evidence supports it.
Summary
Vendor AI governance rests on one fact about liability and one about visibility: the deploying organisation answers for the systems it uses, and the changes that matter most reach it without a procurement event to mark them.
The Decision Rule: Tier by AI Risk, Then Sign Through a Shared Veto
The decision rule that runs through the whole programme is to classify the vendor by what its AI does, and to let that classification set the tier before spend or contract value gets a say. The overlay on the existing workflow, the role assignments and the veto arrangement all follow from that one rule. Whatever tier the classification sets, the vendor’s evidence gets read against the adequacy bar for that tier, and the people who hold the vetoes decide whether it proceeds.
Sequencing matters as much as the rule. Classification comes first because it determines how much evidence is enough. Evidence comes second, judged for whether it predicts behaviour on the organisation’s own cases and not for whether a document exists. The approve, condition or reject decision comes third, made by the review board with conditions carrying expiry dates. Concentration then feeds the exit plan without overturning the verdict.
Applied consistently, the rule also protects the organisation from its own habits. Procurement pressure, familiarity with a large supplier and the comfort of a certificate all push toward the old sort key, and a written rule with named vetoes gives reviewers a way to resist without a personal confrontation. The first cycle on three vendors shows where the rule is missing from current practice, and the team can decide which gap to close next.
The Failure Mode: Measuring Effort and Calling It Detection
The failure mode that separates a working programme from a decorative one is measuring effort and reporting it as detection. Coverage counts, completed questionnaires and certificates on file all describe activity. They say nothing about whether the programme would notice a vendor’s model changing, a bias finding going unfixed or an AI feature appearing in a tool nobody classified.
The guard against this is to ask of every measure which events it would catch. Coverage counts, completed questionnaires and certificates on file would catch none of them, because they record what the team did and not what the programme can see. A measure earns its place on the scorecard when a change in the vendor’s model or a new AI feature would move it.
The same discipline applies to monitoring. A calendar produces a comforting rhythm of reviews that the vendor’s release cycle ignores, so the programme’s real coverage is the set of events it would catch, and a review date only says when it last looked. Teams that keep the outcome measures in front of the review board tend to notice early when a comfortable number has stopped meaning anything, and the adjustment that follows is theirs to choose.
Related in this cluster
- AI Governance and Responsible AI: The Complete Enterprise Guide
- AI Risk Management and Compliance: Frameworks and Controls
- AI Model Governance and Lifecycle Management
- AI Ethics and Fairness: Principles, Frameworks, and Implementation
- AI Transparency and Explainability: XAI Techniques and Tools
- AI Accountability and Responsibility: Frameworks for Assigning Ownership
Anonymous. Counted, not tracked.
Where is your organisation with this right now?
What is the hardest part where you are?
In a sentence: what are you trying to work out right now?
No names, no company. Anonymous. Counted, not tracked.