Executive summary
AI vendor evaluation is a decision process, not a security questionnaire sent in isolation. The review should connect the intended business outcome to the exact data flow, system behavior, evidence, contract, pilot, owners, and residual risk.
A vendor can be acceptable for public-content drafting and unacceptable for restricted records or high-impact decisions.
“Enterprise-grade” is a claim. An applicable agreement, scoped report, configuration test, or observed control is evidence.
A strong overall score cannot compensate for an unresolved critical issue such as unapproved data reuse or excessive AI authority.
A favorable review normally supports a controlled pilot with conditions—not unrestricted production deployment.
Fastest practical route: use the free AI Vendor Risk Assessment to capture the context, answer the 28 control questions, identify evidence gaps, and generate a decision brief.
1. Define the decision before reviewing the vendor
A review without a defined use case usually produces vague answers. Before sending any questionnaire, write a one-page scope. Identify what the system will do, who will use it, what decisions it may influence, which data it can receive, what it can retrieve, and whether it can act in another system.
Describe the complete system
The “vendor” is rarely the whole system. A deployed AI workflow can include the vendor application, a foundation model, retrieval sources, browser extensions, plugins, APIs, orchestration tools, identity services, logging platforms, subprocessors, and human reviewers. Draw the path from input to output and action.
State the job to be done, affected users, accountable owner, prohibited uses, and success criteria.
Classify prompts, files, retrieved content, outputs, feedback, logs, metadata, and connected-system data.
Distinguish drafting, recommendation, human-approved action, bounded action, and autonomous action.
Record plan, region, hosting, tenant model, identity setup, integrations, model choices, and configurations.
Identify safety, rights, financial, employment, education, health, customer, operational, and reputational effects.
Name the business, security, privacy, legal, procurement, technical, accessibility, and risk reviewers required.
Assign an initial risk tier
| Tier | Example context | Typical review depth | Default next step |
|---|---|---|---|
| Low | Optional drafting with public information and meaningful human review | Basic product, data, access, terms, and quality checks | Small time-boxed trial |
| Moderate | Internal information, customer-facing drafts, or workflow recommendations | Cross-functional evidence review plus representative testing | Bounded pilot with monitoring |
| High | Personal or confidential data, material customer effects, or system actions | Specialist review, stronger contract, threat testing, fallback, and approval gates | Pause until blockers close |
| Critical | Restricted data, safety-critical use, high-impact decisions, or broad autonomous authority | Formal governance, legal analysis, deep testing, independent evidence, and senior risk acceptance | Redesign or tightly controlled evaluation |
Risk tier is use-case specific. Do not copy a vendor approval from one department to a different workflow without checking data, impact, integrations, region, plan, and authority again.
2. Build an evidence standard
Vendor due diligence fails when every answer is treated as equally reliable. Define what “Verified” means before scoring the vendor.
Useful for discovery, but usually not enough for approval. Record the claim and request the underlying control or commitment.
Product docs, trust-center content, privacy notices, system cards, limitations, and architecture descriptions.
A signed DPA, order form, SLA, security schedule, data-use term, incident notice, or negotiated protection.
Current scoped assurance, configuration verification, technical test, pilot evidence, audit artifact, or observed behavior.
Apply five evidence tests
- Current: Is the evidence recent enough for the vendor’s current system and control environment?
- Applicable: Does it cover the exact product, plan, model, region, service, and entity you will use?
- Specific: Does it state a testable control, scope, limitation, exception, or commitment?
- Independent where needed: Is material assurance produced or reviewed by a competent party separate from the sales claim?
- Repeatable: Can the control be monitored, tested, reported, or enforced after the purchase?
| Status | Meaning | How to use it |
|---|---|---|
| VERIFIED | Current, applicable evidence supports the control for this decision. | Record source, scope, reviewer, date, exceptions, and next review. |
| PARTIAL | Some evidence exists, but scope, implementation, testing, or commitment is incomplete. | Create a condition with an owner and due date. |
| UNKNOWN | The control has not been established. | Treat it as an evidence gap, not an assumed control. |
| NO | The control is absent, contradicted, or unacceptable for the use case. | Mitigate, redesign, negotiate, or stop. |
3. Review privacy, data use, and data protection
Begin with a product-specific data-flow diagram. Include prompt and file content, retrieved documents, outputs, user feedback, conversation history, embeddings, account metadata, administrative logs, support access, abuse monitoring, and telemetry. Identify each organization that can receive or process the data.
Questions that require precise answers
- What data is collected at the user, account, workspace, API, connector, support, and model layers?
- Is customer content used for training, fine-tuning, evaluation, human review, safety improvement, or another secondary purpose?
- Is the relevant setting opt-in, opt-out, contractually disabled, or controlled only through product configuration?
- What are the retention periods for prompts, files, outputs, deleted content, backups, logs, abuse records, and support cases?
- Can administrators enforce retention and delete content? How is deletion propagated to copies and subprocessors?
- Where is data stored and processed, and which subprocessors, model providers, or support locations are involved?
- Are there plan-specific differences in training, isolation, retention, residency, support access, or controls?
- How are access, correction, deletion, export, restriction, objection, and other applicable rights supported?
Confirm the complete data chain
A “no training on your data” statement does not answer every privacy question. Content may still be retained for service delivery, logged for safety, exposed to support personnel, processed by another model provider, copied into a connector, or included in telemetry. Ask separately about purpose, access, retention, location, onward transfer, deletion, and contractual enforceability.
Blocker: sensitive or restricted data is in scope, but the organization cannot verify applicable data-use terms, retention, deletion, access, subprocessors, or the data processing agreement.
Map legal and regulatory roles
Identify the applicable jurisdictions and whether the parties act as controller, processor, provider, deployer, developer, distributor, or another regulated role. Requirements vary by use case and location. For EU-related activity, confirm the current AI Act classification and obligations on the decision date rather than relying on a static sales statement. Regulatory review should be completed by qualified counsel or the appropriate internal specialist.
4. Review conventional and AI-specific security
Security evidence should cover the whole service life cycle: secure design, development, deployment, operation, change, incident handling, and exit.
Identity, access, and tenant control
- Require appropriate SSO, MFA, role-based access, least privilege, user lifecycle controls, service-account governance, and periodic access review.
- Verify administrator capabilities, support access, privileged activity logging, export restrictions, sharing defaults, and workspace discovery.
- Test tenant separation and authorization for files, conversations, agents, connectors, knowledge stores, APIs, and administrative functions.
Encryption, secrets, and infrastructure
- Confirm encryption in transit and at rest, key-management responsibilities, backup protection, and any customer-managed key option required by policy.
- Check how API keys, OAuth tokens, connector credentials, system prompts, model endpoints, and service secrets are stored, rotated, scoped, and revoked.
- Understand hosting boundaries, production access, environment separation, asset inventories, vulnerability management, patching, and secure configuration.
Secure development and supply chain
Ask how the vendor inventories code, models, datasets, prompts, components, plugins, and dependencies. Review secure development practices, code and dependency scanning, change controls, testing, provenance, vulnerability disclosure, penetration testing, remediation, and supply-chain risk. A conventional assurance report may support this review, but verify its product scope, time period, exceptions, complementary user controls, and relevance to the AI service.
Detection, response, and resilience
- What events are logged, who can access them, how long they are retained, and can customers export them?
- How does the vendor detect abuse, data leakage, unauthorized access, harmful outputs, model or prompt changes, and abnormal consumption?
- What incident-notification trigger, content, channel, and timeline apply under the signed agreement?
- What are the recovery objectives, tested continuity procedures, dependency risks, status communications, and customer fallback options?
The NCSC secure AI guidance organizes work across secure design, secure development, secure deployment, and secure operation and maintenance. Use those stages to find gaps that a point-in-time questionnaire can miss.
5. Review AI-specific behavior and governance
AI systems introduce uncertainty and attack paths beyond conventional SaaS. Review the complete application, not only the base model. Retrieval, tools, memory, system prompts, plugins, agent loops, and human workflow can create or reduce risk.
| Risk area | What to ask | What to test | Common control |
|---|---|---|---|
| Prompt injection | How are untrusted instructions separated from trusted policy and tools? | Direct and indirect injection through files, web content, email, RAG, and connectors | Input boundaries, least privilege, tool allowlists, approval gates, monitoring |
| Sensitive disclosure | How are secrets, personal data, tenant data, prompts, memory, and retrieved content protected? | Cross-user access, extraction attempts, memorization, logs, exports, support paths | Data minimization, access control, isolation, filtering, redaction, testing |
| Improper output handling | Can model output reach code, queries, browsers, messages, or business systems? | Injection, unsafe rendering, command generation, invalid structured output | Validation, sanitization, typed schemas, sandboxing, human approval |
| Excessive agency | What actions, resources, permissions, spend, and duration can the AI control? | Chained actions, privilege escalation, goal drift, loops, unexpected side effects | Minimum permissions, bounded tasks, rate and spend limits, kill switch |
| Misinformation | What limitations, grounding, confidence, citations, and evaluation evidence exist? | Representative facts, edge cases, conflicts, freshness, adversarial examples | Grounding, source display, abstention, human review, fallback |
| Supply chain and change | Which models, data, libraries, endpoints, and subprocessors can change? | Version changes, degraded outputs, new permissions, dependency failure | Inventory, version pinning where possible, change notice, regression tests |
| Unbounded consumption | How are tokens, calls, jobs, retries, storage, and tool actions constrained? | Large input, recursion, retry storms, denial of wallet, resource exhaustion | Quotas, timeouts, circuit breakers, budgets, anomaly alerts |
Demand meaningful system documentation
Request intended use, excluded use, architecture, model and component inventory, data sources where appropriate, evaluation methods, known limitations, failure modes, safety controls, monitoring, change practices, and escalation routes. Documentation should help your team predict behavior and operate the system—not merely describe its benefits.
Evaluate with your data and workflow
Generic benchmarks cannot prove fitness for your use case. Build representative test sets that include normal work, difficult cases, sensitive inputs, multilingual content where relevant, adversarial instructions, incomplete data, retrieval conflicts, and operational failures. Measure the consequence of errors, not only average output quality.
Human review is a designed control only when it is realistic. Reviewers need the time, information, authority, training, interface, escalation path, and fallback required to detect and correct the expected failures.
6. Convert important controls into contract terms
Product features can change. Contract review determines which promises apply to your plan and what happens when a control, model, subprocessor, price, or service changes. Coordinate the technical review with qualified legal and procurement specialists.
| Contract area | Questions to resolve | Why it matters |
|---|---|---|
| Scope and order of precedence | Which service, plan, region, features, policies, URLs, and documents form the agreement? | Prevents a favorable document from being overridden or excluded. |
| Data use and DPA | What processing instructions, purposes, roles, retention, deletion, transfers, and subprocessors apply? | Turns data expectations into applicable commitments. |
| Security schedule | Which controls, assurance, testing, remediation, access, encryption, and continuity duties are binding? | Aligns the promised security baseline with the actual use case. |
| Incident notice | What triggers notice, how quickly, with what content, updates, cooperation, and evidence? | Supports your own response and notification duties. |
| AI-specific change | Can models, capabilities, data practices, limitations, or safety controls change? What notice and options exist? | Prevents material risk from changing silently. |
| IP and content | Who owns inputs, outputs, customizations, prompts, feedback, and derived artifacts? What claims or restrictions apply? | Clarifies permitted business use and dispute handling. |
| Audit and information | What reports, evidence, testing summaries, questionnaires, audit rights, or regulator cooperation are available? | Keeps verification possible after signing. |
| SLA and support | What availability, support, performance, change, remedy, and escalation commitments apply? | Connects operational dependency to enforceable service expectations. |
| Liability and indemnity | How are confidentiality, security, privacy, IP, regulatory, and AI-related losses allocated? | Aligns risk allocation with the likely impact and bargaining context. |
| Exit and portability | Can you export data, configuration, logs, prompts, and knowledge? What is deleted and when? | Reduces lock-in and enables a safe transition or shutdown. |
Do not rely on a screenshot of a public policy. Record the version reviewed and confirm whether the signed agreement incorporates it, allows unilateral change, gives notice, and provides an acceptable remedy or exit.
7. Run a bounded, representative pilot
A document review tells you what should happen. A pilot shows what happens in your environment. Keep the scope small enough to stop, observe, and reverse.
- Define the baseline: measure the current process, quality, effort, delay, cost, error, exception, and incident rates before introducing AI.
- Choose representative cases: include routine, difficult, incomplete, sensitive, multilingual, adversarial, and failure scenarios in realistic proportions.
- Constrain access and authority: use the minimum data, users, connectors, permissions, actions, time, and spending required to test the hypothesis.
- Define review and fallback: state who checks outputs, what evidence they see, when they can override, and how work continues if the AI fails.
- Set thresholds before results: establish minimum quality, maximum serious error, review effort, cost, reliability, and safety conditions in advance.
- Log cases and exceptions: capture inputs by category, outputs, corrections, review time, failures, incidents, escalation, fallback, and outcome.
- Test abuse and change: probe injection, disclosure, unsafe output, excessive permissions, model change, outage, connector failure, and spend limits.
- Make an accountable decision: approve, condition, pause, or reject with named owners, unresolved risks, evidence, and a review date.
For a detailed measurement framework, use How to Measure an AI Automation Pilot in 2026 with the AI Automation Pilot Tracker and Pilot Scorecard.
8. Score the evidence—and keep decision gates
A useful scoring model makes uncertainty visible. Score privacy and data, security, AI governance, contract and accountability, and operational resilience separately. Weight them for your organization, but never let the weighted average erase a critical blocker.
Weighted maturity = Σ(answer value × question importance) ÷ Σ(maximum value × question importance)Verified = 100 · Partial = 55 · Unknown = 15 · No = 0Residual planning risk = 100 − control maturity + context uplift + weak-evidence upliftThe numbers create consistency; they do not create certainty. A score is not a probability of breach, a compliance determination, an independent certification, or permission to deploy.
Proceed to bounded pilot
Applicable evidence is strong, no blocker is open, and representative testing can occur within defined limits.
Proceed with controls
Gaps are manageable through explicit conditions, owners, dates, monitoring, and pilot limits.
Pause and investigate
Material unknowns, weak evidence, or a critical gate require resolution before the next step.
Do not approve yet
The use case or control set is unacceptable. Redesign, choose another vendor, or obtain materially stronger protection.
9. AI vendor red flags
A red flag is not always an automatic rejection. It is a signal that the current evidence or design cannot support the intended decision.
Scope ambiguity
The vendor cannot confirm which product, model, plan, region, entity, or subprocessor the evidence covers.
Data-use ambiguity
Training, evaluation, human review, retention, deletion, or secondary-use answers are vague or contradictory.
No meaningful limitations
Documentation describes benefits but not failure modes, excluded uses, evaluation limits, or safe operating conditions.
Excessive default access
The service requests broad connectors, persistent credentials, actions, sharing, or data access beyond the use case.
Weak change control
Models, terms, data practices, subprocessors, controls, or capabilities can change without useful notice or recourse.
No workable exit
The organization cannot export required artifacts, revoke access, transition service, confirm deletion, or preserve evidence.
Marketing as evidence
“Secure,” “private,” “compliant,” or “responsible” claims are not supported by applicable documents or tests.
Incident opacity
Notification triggers, timing, communication, investigation support, or past-event explanations are inadequate.
Unrealistic human oversight
The workflow says “human in the loop,” but reviewers lack time, authority, information, training, or fallback.
10. Copy-ready AI vendor questionnaire
Send the questionnaire with a short scope statement describing your planned product, plan, region, data, integrations, authority, and decision stage. Ask the vendor to link evidence and mark anything that does not apply.
AI vendor due-diligence request
Procurement tip: ask the vendor to keep the original numbering. That makes gaps, follow-ups, evidence, contract changes, and reassessment easier to trace.
11. Create an accountable decision record
The final output should be short enough to review and specific enough to defend. Preserve the scope and evidence that existed at the time of the decision.
Minimum decision record
Set reassessment triggers
Reassess after a material model, feature, agent, connector, subprocessor, hosting, region, data-use, contract, security, ownership, or pricing change; after a significant incident or unexplained degradation; before expanding data sensitivity, user population, business impact, or AI authority; and at a risk-based recurring interval.
12. A practical 30-day review workflow
| Period | Work | Primary output | Decision gate |
|---|---|---|---|
| Days 1–3 | Define use case, system boundary, data, authority, impact, owners, prohibited use, and initial tier. | One-page scope and data flow | Is the use case suitable for review? |
| Days 4–10 | Collect privacy, security, AI, contract, architecture, assurance, subprocessor, and continuity evidence. | Evidence register and open questions | Are any blockers already visible? |
| Days 11–15 | Run specialist reviews, vendor follow-up, architecture threat review, and contract issue identification. | Dimension findings and conditions | Can a bounded pilot be designed safely? |
| Days 16–25 | Run representative functional, quality, security, abuse, failure, review, and fallback tests. | Case log and pilot evidence | Were thresholds and stop conditions met? |
| Days 26–30 | Close material gaps, finalize contract conditions, record residual risk, owners, monitoring, and reassessment. | Decision brief and approval record | Proceed, condition, pause, or reject? |
Common evaluation mistakes
- Beginning with a generic questionnaire instead of a defined use case and data flow.
- Treating a certification, assurance report, trust page, or benchmark as complete approval.
- Reviewing the foundation model while ignoring the application, retrieval, tools, agents, and human workflow.
- Marking unanswered questions as acceptable because the vendor is well known.
- Testing only successful demos instead of difficult, adversarial, sensitive, and failure cases.
- Approving “human review” without measuring whether reviewers can detect and correct errors.
- Negotiating the contract after the technical team has already committed to the vendor.
- Failing to document conditions, owners, due dates, monitoring, change triggers, and exit.
Frequently asked questions
Methodology and primary sources
This guide synthesizes practical vendor review steps from established risk, security, and AI governance resources. Standards and laws evolve; verify the current version and applicability when making a decision.
- NIST AI Risk Management Framework — the Govern, Map, Measure, and Manage functions and trustworthy AI characteristics.
- NIST Generative AI Profile — a cross-sector companion to the AI RMF for generative AI risks.
- NIST AI RMF Playbook — suggested actions for operationalizing framework outcomes.
- NCSC Guidelines for Secure AI System Development — secure design, development, deployment, operation, and maintenance.
- OWASP Top 10 for LLM and GenAI Applications — prompt injection, sensitive disclosure, supply chain, improper output handling, excessive agency, and related risks.
- European Commission AI Act overview — current official information on roles, obligations, implementation, and supporting guidance.
Important: this article is general educational material. It is not legal advice, a compliance determination, an independent vendor certification, or authorization to purchase or deploy an AI system.