2026 AI Vendor Due-Diligence Guide

How to Evaluate an AI Vendor in 2026: Security, Privacy, Risk & Due Diligence

A practical, evidence-based process for moving from vendor claims to an accountable shortlist, bounded pilot, contract decision, and ongoing review.

Executive summary

AI vendor evaluation is a decision process, not a security questionnaire sent in isolation. The review should connect the intended business outcome to the exact data flow, system behavior, evidence, contract, pilot, owners, and residual risk.

Start with context

A vendor can be acceptable for public-content drafting and unacceptable for restricted records or high-impact decisions.

Separate claims from evidence

“Enterprise-grade” is a claim. An applicable agreement, scoped report, configuration test, or observed control is evidence.

Do not average away blockers

A strong overall score cannot compensate for an unresolved critical issue such as unapproved data reuse or excessive AI authority.

Approve a bounded next step

A favorable review normally supports a controlled pilot with conditions—not unrestricted production deployment.

Fastest practical route: use the free AI Vendor Risk Assessment to capture the context, answer the 28 control questions, identify evidence gaps, and generate a decision brief.

1. Define the decision before reviewing the vendor

A review without a defined use case usually produces vague answers. Before sending any questionnaire, write a one-page scope. Identify what the system will do, who will use it, what decisions it may influence, which data it can receive, what it can retrieve, and whether it can act in another system.

Describe the complete system

The “vendor” is rarely the whole system. A deployed AI workflow can include the vendor application, a foundation model, retrieval sources, browser extensions, plugins, APIs, orchestration tools, identity services, logging platforms, subprocessors, and human reviewers. Draw the path from input to output and action.

Outcome and users

State the job to be done, affected users, accountable owner, prohibited uses, and success criteria.

Data

Classify prompts, files, retrieved content, outputs, feedback, logs, metadata, and connected-system data.

Authority

Distinguish drafting, recommendation, human-approved action, bounded action, and autonomous action.

Deployment

Record plan, region, hosting, tenant model, identity setup, integrations, model choices, and configurations.

Impact

Identify safety, rights, financial, employment, education, health, customer, operational, and reputational effects.

Accountability

Name the business, security, privacy, legal, procurement, technical, accessibility, and risk reviewers required.

Assign an initial risk tier

TierExample contextTypical review depthDefault next step
LowOptional drafting with public information and meaningful human reviewBasic product, data, access, terms, and quality checksSmall time-boxed trial
ModerateInternal information, customer-facing drafts, or workflow recommendationsCross-functional evidence review plus representative testingBounded pilot with monitoring
HighPersonal or confidential data, material customer effects, or system actionsSpecialist review, stronger contract, threat testing, fallback, and approval gatesPause until blockers close
CriticalRestricted data, safety-critical use, high-impact decisions, or broad autonomous authorityFormal governance, legal analysis, deep testing, independent evidence, and senior risk acceptanceRedesign or tightly controlled evaluation

Risk tier is use-case specific. Do not copy a vendor approval from one department to a different workflow without checking data, impact, integrations, region, plan, and authority again.

2. Build an evidence standard

Vendor due diligence fails when every answer is treated as equally reliable. Define what “Verified” means before scoring the vendor.

Level 1Marketing claim

Useful for discovery, but usually not enough for approval. Record the claim and request the underlying control or commitment.

Level 2Published documentation

Product docs, trust-center content, privacy notices, system cards, limitations, and architecture descriptions.

Level 3Applicable commitment

A signed DPA, order form, SLA, security schedule, data-use term, incident notice, or negotiated protection.

Level 4Tested or assured

Current scoped assurance, configuration verification, technical test, pilot evidence, audit artifact, or observed behavior.

Apply five evidence tests

  1. Current: Is the evidence recent enough for the vendor’s current system and control environment?
  2. Applicable: Does it cover the exact product, plan, model, region, service, and entity you will use?
  3. Specific: Does it state a testable control, scope, limitation, exception, or commitment?
  4. Independent where needed: Is material assurance produced or reviewed by a competent party separate from the sales claim?
  5. Repeatable: Can the control be monitored, tested, reported, or enforced after the purchase?
StatusMeaningHow to use it
VERIFIEDCurrent, applicable evidence supports the control for this decision.Record source, scope, reviewer, date, exceptions, and next review.
PARTIALSome evidence exists, but scope, implementation, testing, or commitment is incomplete.Create a condition with an owner and due date.
UNKNOWNThe control has not been established.Treat it as an evidence gap, not an assumed control.
NOThe control is absent, contradicted, or unacceptable for the use case.Mitigate, redesign, negotiate, or stop.

3. Review privacy, data use, and data protection

Begin with a product-specific data-flow diagram. Include prompt and file content, retrieved documents, outputs, user feedback, conversation history, embeddings, account metadata, administrative logs, support access, abuse monitoring, and telemetry. Identify each organization that can receive or process the data.

Questions that require precise answers

  • What data is collected at the user, account, workspace, API, connector, support, and model layers?
  • Is customer content used for training, fine-tuning, evaluation, human review, safety improvement, or another secondary purpose?
  • Is the relevant setting opt-in, opt-out, contractually disabled, or controlled only through product configuration?
  • What are the retention periods for prompts, files, outputs, deleted content, backups, logs, abuse records, and support cases?
  • Can administrators enforce retention and delete content? How is deletion propagated to copies and subprocessors?
  • Where is data stored and processed, and which subprocessors, model providers, or support locations are involved?
  • Are there plan-specific differences in training, isolation, retention, residency, support access, or controls?
  • How are access, correction, deletion, export, restriction, objection, and other applicable rights supported?

Confirm the complete data chain

A “no training on your data” statement does not answer every privacy question. Content may still be retained for service delivery, logged for safety, exposed to support personnel, processed by another model provider, copied into a connector, or included in telemetry. Ask separately about purpose, access, retention, location, onward transfer, deletion, and contractual enforceability.

Blocker: sensitive or restricted data is in scope, but the organization cannot verify applicable data-use terms, retention, deletion, access, subprocessors, or the data processing agreement.

Map legal and regulatory roles

Identify the applicable jurisdictions and whether the parties act as controller, processor, provider, deployer, developer, distributor, or another regulated role. Requirements vary by use case and location. For EU-related activity, confirm the current AI Act classification and obligations on the decision date rather than relying on a static sales statement. Regulatory review should be completed by qualified counsel or the appropriate internal specialist.

4. Review conventional and AI-specific security

Security evidence should cover the whole service life cycle: secure design, development, deployment, operation, change, incident handling, and exit.

Identity, access, and tenant control

  • Require appropriate SSO, MFA, role-based access, least privilege, user lifecycle controls, service-account governance, and periodic access review.
  • Verify administrator capabilities, support access, privileged activity logging, export restrictions, sharing defaults, and workspace discovery.
  • Test tenant separation and authorization for files, conversations, agents, connectors, knowledge stores, APIs, and administrative functions.

Encryption, secrets, and infrastructure

  • Confirm encryption in transit and at rest, key-management responsibilities, backup protection, and any customer-managed key option required by policy.
  • Check how API keys, OAuth tokens, connector credentials, system prompts, model endpoints, and service secrets are stored, rotated, scoped, and revoked.
  • Understand hosting boundaries, production access, environment separation, asset inventories, vulnerability management, patching, and secure configuration.

Secure development and supply chain

Ask how the vendor inventories code, models, datasets, prompts, components, plugins, and dependencies. Review secure development practices, code and dependency scanning, change controls, testing, provenance, vulnerability disclosure, penetration testing, remediation, and supply-chain risk. A conventional assurance report may support this review, but verify its product scope, time period, exceptions, complementary user controls, and relevance to the AI service.

Detection, response, and resilience

  • What events are logged, who can access them, how long they are retained, and can customers export them?
  • How does the vendor detect abuse, data leakage, unauthorized access, harmful outputs, model or prompt changes, and abnormal consumption?
  • What incident-notification trigger, content, channel, and timeline apply under the signed agreement?
  • What are the recovery objectives, tested continuity procedures, dependency risks, status communications, and customer fallback options?

The NCSC secure AI guidance organizes work across secure design, secure development, secure deployment, and secure operation and maintenance. Use those stages to find gaps that a point-in-time questionnaire can miss.

5. Review AI-specific behavior and governance

AI systems introduce uncertainty and attack paths beyond conventional SaaS. Review the complete application, not only the base model. Retrieval, tools, memory, system prompts, plugins, agent loops, and human workflow can create or reduce risk.

Risk areaWhat to askWhat to testCommon control
Prompt injectionHow are untrusted instructions separated from trusted policy and tools?Direct and indirect injection through files, web content, email, RAG, and connectorsInput boundaries, least privilege, tool allowlists, approval gates, monitoring
Sensitive disclosureHow are secrets, personal data, tenant data, prompts, memory, and retrieved content protected?Cross-user access, extraction attempts, memorization, logs, exports, support pathsData minimization, access control, isolation, filtering, redaction, testing
Improper output handlingCan model output reach code, queries, browsers, messages, or business systems?Injection, unsafe rendering, command generation, invalid structured outputValidation, sanitization, typed schemas, sandboxing, human approval
Excessive agencyWhat actions, resources, permissions, spend, and duration can the AI control?Chained actions, privilege escalation, goal drift, loops, unexpected side effectsMinimum permissions, bounded tasks, rate and spend limits, kill switch
MisinformationWhat limitations, grounding, confidence, citations, and evaluation evidence exist?Representative facts, edge cases, conflicts, freshness, adversarial examplesGrounding, source display, abstention, human review, fallback
Supply chain and changeWhich models, data, libraries, endpoints, and subprocessors can change?Version changes, degraded outputs, new permissions, dependency failureInventory, version pinning where possible, change notice, regression tests
Unbounded consumptionHow are tokens, calls, jobs, retries, storage, and tool actions constrained?Large input, recursion, retry storms, denial of wallet, resource exhaustionQuotas, timeouts, circuit breakers, budgets, anomaly alerts

Demand meaningful system documentation

Request intended use, excluded use, architecture, model and component inventory, data sources where appropriate, evaluation methods, known limitations, failure modes, safety controls, monitoring, change practices, and escalation routes. Documentation should help your team predict behavior and operate the system—not merely describe its benefits.

Evaluate with your data and workflow

Generic benchmarks cannot prove fitness for your use case. Build representative test sets that include normal work, difficult cases, sensitive inputs, multilingual content where relevant, adversarial instructions, incomplete data, retrieval conflicts, and operational failures. Measure the consequence of errors, not only average output quality.

Human review is a designed control only when it is realistic. Reviewers need the time, information, authority, training, interface, escalation path, and fallback required to detect and correct the expected failures.

6. Convert important controls into contract terms

Product features can change. Contract review determines which promises apply to your plan and what happens when a control, model, subprocessor, price, or service changes. Coordinate the technical review with qualified legal and procurement specialists.

Contract areaQuestions to resolveWhy it matters
Scope and order of precedenceWhich service, plan, region, features, policies, URLs, and documents form the agreement?Prevents a favorable document from being overridden or excluded.
Data use and DPAWhat processing instructions, purposes, roles, retention, deletion, transfers, and subprocessors apply?Turns data expectations into applicable commitments.
Security scheduleWhich controls, assurance, testing, remediation, access, encryption, and continuity duties are binding?Aligns the promised security baseline with the actual use case.
Incident noticeWhat triggers notice, how quickly, with what content, updates, cooperation, and evidence?Supports your own response and notification duties.
AI-specific changeCan models, capabilities, data practices, limitations, or safety controls change? What notice and options exist?Prevents material risk from changing silently.
IP and contentWho owns inputs, outputs, customizations, prompts, feedback, and derived artifacts? What claims or restrictions apply?Clarifies permitted business use and dispute handling.
Audit and informationWhat reports, evidence, testing summaries, questionnaires, audit rights, or regulator cooperation are available?Keeps verification possible after signing.
SLA and supportWhat availability, support, performance, change, remedy, and escalation commitments apply?Connects operational dependency to enforceable service expectations.
Liability and indemnityHow are confidentiality, security, privacy, IP, regulatory, and AI-related losses allocated?Aligns risk allocation with the likely impact and bargaining context.
Exit and portabilityCan you export data, configuration, logs, prompts, and knowledge? What is deleted and when?Reduces lock-in and enables a safe transition or shutdown.

Do not rely on a screenshot of a public policy. Record the version reviewed and confirm whether the signed agreement incorporates it, allows unilateral change, gives notice, and provides an acceptable remedy or exit.

7. Run a bounded, representative pilot

A document review tells you what should happen. A pilot shows what happens in your environment. Keep the scope small enough to stop, observe, and reverse.

  1. Define the baseline: measure the current process, quality, effort, delay, cost, error, exception, and incident rates before introducing AI.
  2. Choose representative cases: include routine, difficult, incomplete, sensitive, multilingual, adversarial, and failure scenarios in realistic proportions.
  3. Constrain access and authority: use the minimum data, users, connectors, permissions, actions, time, and spending required to test the hypothesis.
  4. Define review and fallback: state who checks outputs, what evidence they see, when they can override, and how work continues if the AI fails.
  5. Set thresholds before results: establish minimum quality, maximum serious error, review effort, cost, reliability, and safety conditions in advance.
  6. Log cases and exceptions: capture inputs by category, outputs, corrections, review time, failures, incidents, escalation, fallback, and outcome.
  7. Test abuse and change: probe injection, disclosure, unsafe output, excessive permissions, model change, outage, connector failure, and spend limits.
  8. Make an accountable decision: approve, condition, pause, or reject with named owners, unresolved risks, evidence, and a review date.

For a detailed measurement framework, use How to Measure an AI Automation Pilot in 2026 with the AI Automation Pilot Tracker and Pilot Scorecard.

8. Score the evidence—and keep decision gates

A useful scoring model makes uncertainty visible. Score privacy and data, security, AI governance, contract and accountability, and operational resilience separately. Weight them for your organization, but never let the weighted average erase a critical blocker.

Control maturityWeighted maturity = Σ(answer value × question importance) ÷ Σ(maximum value × question importance)
Example answer valuesVerified = 100 · Partial = 55 · Unknown = 15 · No = 0
Planning riskResidual planning risk = 100 − control maturity + context uplift + weak-evidence uplift

The numbers create consistency; they do not create certainty. A score is not a probability of breach, a compliance determination, an independent certification, or permission to deploy.

Proceed to bounded pilot

Applicable evidence is strong, no blocker is open, and representative testing can occur within defined limits.

Proceed with controls

Gaps are manageable through explicit conditions, owners, dates, monitoring, and pilot limits.

Pause and investigate

Material unknowns, weak evidence, or a critical gate require resolution before the next step.

Do not approve yet

The use case or control set is unacceptable. Redesign, choose another vendor, or obtain materially stronger protection.

9. AI vendor red flags

A red flag is not always an automatic rejection. It is a signal that the current evidence or design cannot support the intended decision.

Scope ambiguity

The vendor cannot confirm which product, model, plan, region, entity, or subprocessor the evidence covers.

Data-use ambiguity

Training, evaluation, human review, retention, deletion, or secondary-use answers are vague or contradictory.

No meaningful limitations

Documentation describes benefits but not failure modes, excluded uses, evaluation limits, or safe operating conditions.

Excessive default access

The service requests broad connectors, persistent credentials, actions, sharing, or data access beyond the use case.

Weak change control

Models, terms, data practices, subprocessors, controls, or capabilities can change without useful notice or recourse.

No workable exit

The organization cannot export required artifacts, revoke access, transition service, confirm deletion, or preserve evidence.

Marketing as evidence

“Secure,” “private,” “compliant,” or “responsible” claims are not supported by applicable documents or tests.

Incident opacity

Notification triggers, timing, communication, investigation support, or past-event explanations are inadequate.

Unrealistic human oversight

The workflow says “human in the loop,” but reviewers lack time, authority, information, training, or fallback.

10. Copy-ready AI vendor questionnaire

Send the questionnaire with a short scope statement describing your planned product, plan, region, data, integrations, authority, and decision stage. Ask the vendor to link evidence and mark anything that does not apply.

AI vendor due-diligence request

AI VENDOR DUE-DILIGENCE REQUEST Scope 1. Confirm the legal entity, product, plan, deployment model, region, model(s), and service components covered by your answers. 2. List material subprocessors, model providers, hosting providers, data locations, and support locations applicable to this service. 3. Describe the intended uses, excluded uses, known limitations, and material failure modes. Data and privacy 4. Describe every category of customer content, metadata, telemetry, log, feedback, and support data collected or generated. 5. State whether customer content is used for training, fine-tuning, evaluation, human review, safety improvement, or another secondary purpose. Explain defaults and controls. 6. Provide retention and deletion periods for content, history, files, outputs, embeddings, logs, backups, abuse records, and support cases. 7. Explain data residency, cross-border processing, subprocessor changes, and customer notice or objection options. 8. Provide the applicable privacy notice, DPA, transfer mechanism, and process for supporting data-subject requests. Security 9. Describe SSO, MFA, role-based access, user lifecycle, privileged access, support access, and audit logging. 10. Describe tenant isolation, encryption, key management, secrets handling, vulnerability management, penetration testing, and remediation. 11. Identify current assurance evidence and its product, location, control, and time-period scope. Include material exceptions and complementary customer controls. 12. Describe secure development, supply-chain controls, asset inventories, change management, and vulnerability disclosure. 13. Describe incident detection, customer notification triggers and timing, investigation support, continuity, recovery objectives, and tested fallback. AI governance and behavior 14. Provide model or system documentation covering architecture, versions, evaluations, limitations, safety controls, monitoring, and change practices. 15. Explain controls and tests for prompt injection, sensitive information disclosure, improper output handling, excessive agency, misinformation, poisoning, and unbounded consumption. 16. Describe grounding, citations, abstention, confidence, human review, override, rollback, and kill-switch capabilities. 17. Explain how connectors, tools, memory, retrieval sources, actions, rate limits, spend limits, and permissions are controlled and logged. 18. State how customers are notified of material model, capability, safety, data-practice, or performance changes. Contract and operations 19. Identify the documents that form the agreement and their order of precedence. Explain which online terms may change unilaterally. 20. Describe contractual commitments for data use, confidentiality, security, incidents, availability, support, audit information, deletion, and exit. 21. Explain input, output, feedback, prompt, customization, and generated-content ownership and any IP restrictions, warranties, or protections. 22. Provide service dependency, continuity, portability, export, termination-assistance, and verified-deletion details. For each answer, provide a current evidence link or attachment, applicable scope, document date/version, material exception, and named contact for follow-up.

Procurement tip: ask the vendor to keep the original numbering. That makes gaps, follow-ups, evidence, contract changes, and reassessment easier to trace.

11. Create an accountable decision record

The final output should be short enough to review and specific enough to defend. Preserve the scope and evidence that existed at the time of the decision.

Minimum decision record

AI VENDOR DECISION RECORD Vendor / product / plan: Legal entity and region: Business owner: Assessment owner: Decision date and review date: Approved use case: Excluded uses: Users and affected parties: Data categories and sensitivity: Integrations and AI authority: Initial risk tier: Evidence reviewed (document, scope, version/date, reviewer): Privacy and data conclusion: Security conclusion: AI governance and testing conclusion: Contract conclusion: Operational resilience and exit conclusion: Representative pilot results: Critical gates and red flags: Unresolved risks and assumptions: Required controls / contract conditions: Owners and due dates: Fallback, stop, rollback, and incident process: Decision: Proceed to bounded pilot / Proceed with controls / Pause / Do not approve yet Accountable approvers: Reassessment triggers:

Set reassessment triggers

Reassess after a material model, feature, agent, connector, subprocessor, hosting, region, data-use, contract, security, ownership, or pricing change; after a significant incident or unexplained degradation; before expanding data sensitivity, user population, business impact, or AI authority; and at a risk-based recurring interval.

12. A practical 30-day review workflow

PeriodWorkPrimary outputDecision gate
Days 1–3Define use case, system boundary, data, authority, impact, owners, prohibited use, and initial tier.One-page scope and data flowIs the use case suitable for review?
Days 4–10Collect privacy, security, AI, contract, architecture, assurance, subprocessor, and continuity evidence.Evidence register and open questionsAre any blockers already visible?
Days 11–15Run specialist reviews, vendor follow-up, architecture threat review, and contract issue identification.Dimension findings and conditionsCan a bounded pilot be designed safely?
Days 16–25Run representative functional, quality, security, abuse, failure, review, and fallback tests.Case log and pilot evidenceWere thresholds and stop conditions met?
Days 26–30Close material gaps, finalize contract conditions, record residual risk, owners, monitoring, and reassessment.Decision brief and approval recordProceed, condition, pause, or reject?

Common evaluation mistakes

  • Beginning with a generic questionnaire instead of a defined use case and data flow.
  • Treating a certification, assurance report, trust page, or benchmark as complete approval.
  • Reviewing the foundation model while ignoring the application, retrieval, tools, agents, and human workflow.
  • Marking unanswered questions as acceptable because the vendor is well known.
  • Testing only successful demos instead of difficult, adversarial, sensitive, and failure cases.
  • Approving “human review” without measuring whether reviewers can detect and correct errors.
  • Negotiating the contract after the technical team has already committed to the vendor.
  • Failing to document conditions, owners, due dates, monitoring, change triggers, and exit.

Frequently asked questions

Methodology and primary sources

This guide synthesizes practical vendor review steps from established risk, security, and AI governance resources. Standards and laws evolve; verify the current version and applicability when making a decision.

Important: this article is general educational material. It is not legal advice, a compliance determination, an independent vendor certification, or authorization to purchase or deploy an AI system.

Continue your AI vendor evaluation