Executive summary
A useful AI impact assessment connects a specific system and context to affected people, benefit and harm pathways, evidence, controls, residual risk, accountable decisions, and post-deployment learning.
Start before architecture, vendor, contract, or launch choices become expensive to change.
Include models, data, prompts, retrieval, tools, integrations, vendors, people, policies, and operating conditions.
Direct users are not the only stakeholders. Include people subject to outputs, actions, errors, surveillance, delay, or exclusion.
A high average score must not hide a critical safety, rights, privacy, fairness, or human-oversight gap.
Use the companion tool: the free AI Impact Assessment Generator maps eight impact domains, evaluates 29 controls, identifies decision gates, and creates a mitigation and monitoring brief.
What is an AI impact assessment?
An AI impact assessment—or AIA—is a structured process for understanding how an AI-enabled system may affect people, organizations, communities, rights, safety, and operations. It documents the system and purpose, identifies affected people, maps benefits and harms, evaluates likelihood and severity, tests the strength of controls, selects mitigations, records residual risk, and defines monitoring and reassessment.
The assessment should cover the socio-technical system, not only the model. A model can appear accurate in a benchmark while the deployed system fails because of poor data, an unsafe interface, excessive permissions, weak review, automation bias, inaccessible challenge routes, vendor change, or a workflow that creates pressure to accept the output.
What an AIA should answer
What problem is being addressed, why AI is appropriate, what alternatives exist, and what outcome defines success?
Who uses the system, who is subject to it, who may benefit, who may carry burden, and who may be overlooked?
What physical, material, rights, privacy, fairness, dignity, access, operational, or societal impacts are plausible?
Which preventive, detective, corrective, human, technical, contractual, and governance controls are implemented and evidenced?
Does the system proceed, proceed with conditions, pause for redesign, or stop—and who has authority to decide?
How will outcomes, complaints, errors, disparities, incidents, overrides, changes, and new impacts trigger intervention?
Not a universal compliance certificate: different laws, sectors, jurisdictions, and organizations may require specific assessments, records, consultation, publication, review, or approval. A general AIA can coordinate these activities but does not automatically satisfy them.
AI impact assessment vs DPIA, vendor risk, and security review
These reviews overlap, but they answer different questions. Treat them as connected evidence streams rather than substitutes.
| Review | Primary question | Typical scope | Relationship to AIA |
|---|---|---|---|
| AI Impact Assessment | How could this AI system affect people, rights, safety, fairness, operations, and society—and should it proceed? | Purpose, people, benefit, harm, data, quality, fairness, safety, security, oversight, remedy, lifecycle | Provides the broad decision record and links specialist reviews. |
| DPIA / privacy impact assessment | How does personal-data processing affect people's rights and freedoms, and how will those risks be reduced? | Nature, scope, context, purpose, necessity, proportionality, data flow, rights, privacy and security risks | May be a required specialist input when personal data or high-risk processing is involved. |
| Fundamental-rights or equality review | Could the system affect protected rights, equality, access, due process, or vulnerable people? | Legal and social context, groups, decisions, explanation, participation, challenge, remedy, discriminatory effects | Deepens the rights and affected-person sections. |
| Vendor Risk Assessment | Can the supplier, product, evidence, contract, and operations support the required safeguards? | Privacy, security, AI governance, contract, continuity, subprocessors, change, exit | Tests whether a third party can meet controls identified by the AIA. |
| Security threat assessment | How could the system be attacked, misused, compromised, or made unsafe? | Assets, actors, trust boundaries, attack paths, access, injection, supply chain, detection, response, recovery | Provides technical risk and mitigation evidence. |
| Safety or domain validation | Is the system sufficiently safe and fit for its intended domain and operating conditions? | Hazards, failure modes, validation, performance limits, fallback, human factors, monitoring | May be essential in health, transport, infrastructure, or other safety-relevant contexts. |
The correct combination depends on what the system does. For example, an internal writing assistant using public content may need a lighter review than a system that ranks job applicants, recommends medical action, detects fraud, or can execute transactions.
Step 0: Set the trigger, team, and decision owner
Define when an assessment is required and who can change or stop the project. Common triggers include personal or sensitive data, consequential recommendations or decisions, surveillance or profiling, vulnerable groups, safety relevance, public-facing use, broad scale, autonomous action, essential services, new or uncertain technology, or material change to an existing system.
Build a context-specific team
Owns the purpose, benefits, resources, risk decision, operating outcome, and authority to pause or stop.
Understand real cases, exceptions, consequences, workarounds, human judgment, and service conditions.
Reveal impacts, barriers, expectations, context, and remedies that internal teams may not see.
Explain architecture, data, models, tests, limitations, integrations, authority, changes, and monitoring.
Map applicable duties, data roles, threats, contracts, records, escalation, and required specialist reviews.
Challenge inclusion, third-party evidence, residual risk, control ownership, and independent verification.
The team should have enough authority and diversity of experience to challenge the business case. If the assessment can only recommend cosmetic changes after the launch decision is fixed, it is too late.
Step 1: Define the complete AI system and context
A strong assessment begins with a boundary that another reviewer can understand. “We use AI for support” is not a system definition.
- State the purpose and outcome: describe the problem, intended benefit, decision or action, users, affected people, and what success means.
- Describe the workflow: map triggers, inputs, models, prompts, retrieval, rules, tools, outputs, human review, actions, exceptions, and fallback.
- Inventory data: include training, fine-tuning, evaluation, prompt, file, retrieved, output, feedback, metadata, log, and connected-system data.
- Map organizations and components: identify vendors, model providers, subprocessors, APIs, hosting, connectors, datasets, open-source components, and human service providers.
- Define authority: distinguish drafting, recommendation, material determination, bounded action, and autonomous action. Record permissions, limits, and approval points.
- Describe operating context: include scale, frequency, geography, language, accessibility, time pressure, user skill, environmental conditions, and dependency on the service.
- Record exclusions and alternatives: state prohibited use, cases that require a person, non-AI alternatives, and the fallback if the system is unavailable or unsuitable.
Write an assessment scope statement
Example: “This assessment covers the enterprise-plan assistant used by 35 support agents to recommend—not automatically issue—refund outcomes for consumer orders in the Netherlands. It uses order history, policy documents, and agent notes. A trained agent must approve every action. It excludes fraud decisions, account closure, vulnerable-customer cases, and refunds above €500.”
This statement is specific enough to test. If the team later enables automatic refunds, adds a new country, uses more sensitive data, expands to fraud, or changes the model, the assessment boundary has changed.
Step 2: Identify and involve affected people
List everyone who may experience benefit, burden, error, delay, surveillance, exclusion, persuasion, or a changed relationship because of the system. Include direct users, people subject to decisions, people whose data is used, bystanders, frontline workers, communities, and people affected indirectly by resource allocation.
Map power, vulnerability, and access
- Can the person refuse the AI process or use a meaningful alternative without penalty?
- Does the organization control essential employment, education, finance, health, housing, safety, or public services?
- Could disability, language, age, digital access, literacy, immigration status, income, or social position change the impact?
- Can affected people understand the system's role, correct data, challenge an outcome, contact a competent human, and obtain timely remedy?
- Who benefits from speed or cost savings, and who carries additional review work, risk, proof, delay, or emotional burden?
Make participation meaningful
Consultation is not a survey sent after design. Engage early enough to change requirements, scope, data, interface, notices, metrics, safeguards, fallback, or the decision not to deploy. Use accessible methods, compensate participation where appropriate, explain how feedback will be used, and report what changed.
Affected-person interview prompts
Step 3: Map benefits, harms, and impact pathways
Do not begin with a score. Begin with a scenario: cause → system behavior → affected person or group → consequence → duration and reversibility. A clear pathway makes the risk testable and the mitigation assignable.
Unsafe action, delayed care, injury, harmful instruction, unsafe product, environmental damage, or failure during an emergency.
Loss of due process, freedom, lawful treatment, explanation, participation, challenge, remedy, or access to an accountable decision-maker.
Employment, income, credit, insurance, benefits, price, fraud loss, opportunity, contractual position, or financial burden.
Unequal error, quality, allocation, access, representation, treatment, burden, or outcome across relevant people and groups.
Surveillance, inference, re-identification, disclosure, secondary use, unwanted persistence, manipulation, or loss of control over data.
Denial, delay, ranking, reduced quality, exclusion, reduced ability to obtain help, or displacement of a meaningful human channel.
Deception, manipulation, dependency, stigma, chilling effects, distress, loss of agency, or being treated as a score rather than a person.
Outage, fraud, security compromise, misinformation, resource loss, institutional harm, market effects, or community-level consequences.
Map benefits with the same discipline
Describe who receives each benefit, under what conditions, how it will be measured, and what trade-off it creates. “Efficiency” is incomplete. A measurable statement might be: “Reduce median first-response time from six hours to two hours without increasing serious factual errors, escalation failures, accessibility complaints, or agent review time above the agreed threshold.”
Distribution matters: an average benefit can coexist with serious harm to a smaller group. Record subgroup outcomes, rare high-consequence failures, and who carries the cost of correction or appeal.
Step 4: Rate severity and likelihood before controls
Rate each plausible impact before assuming planned safeguards work. This is the inherent impact. Use evidence and a documented scale so different reviewers can understand the rating.
A practical four-level scale
| Level | Severity guide | Likelihood guide |
|---|---|---|
| 1 | Minor, short-lived, easily corrected, narrow, and unlikely to affect important rights or essential needs. | Rare under intended and reasonably foreseeable conditions. |
| 2 | Meaningful inconvenience, delay, loss, distress, or unequal quality that is correctable with effort. | Possible; credible scenarios or limited evidence indicate it could occur. |
| 3 | Serious material, rights, safety, privacy, economic, access, or discriminatory effect; difficult or slow to reverse. | Probable in some operating conditions, groups, failure modes, or repeated use. |
| 4 | Critical, widespread, persistent, irreversible, safety-threatening, rights-threatening, or severe impact on vulnerable people or essential services. | Likely or recurring without substantial design change or effective controls. |
Inherent impact score = severity (1–4) × likelihood (1–4)Severity asks “How bad could the plausible consequence be?” Likelihood asks “How reasonably often could that pathway occur?”| Likelihood ↓ / Severity → | 1 Minor | 2 Moderate | 3 Serious | 4 Critical |
|---|---|---|---|---|
| 4 Likely | 4 | 8 | 12 | 16 |
| 3 Probable | 3 | 6 | 9 | 12 |
| 2 Possible | 2 | 4 | 6 | 8 |
| 1 Rare | 1 | 2 | 3 | 4 |
Document uncertainty
A precise number can hide weak evidence. Record the assumptions, evidence quality, disagreement, confidence, missing groups, untested conditions, and what information would change the rating. Where impact could be severe and evidence is weak, use a conservative decision gate or limit the scope until uncertainty is reduced.
Step 5: Evaluate controls across the lifecycle
For every material impact pathway, identify controls that prevent it, detect it, reduce its consequence, support correction or remedy, and enable stopping. Grade the actual implementation—not the quality of the policy wording.
Participation, notice, explanation, accessibility, meaningful choice, challenge, correction, appeal, remedy, and alternative service.
Purpose, necessity, minimization, provenance, representativeness, quality, access, retention, deletion, rights, and specialist assessment.
Use-case fitness, representative tests, subgroup outcomes, uncertainty, abstention, fallback, reproducibility, and drift monitoring.
Threat modeling, injection testing, least privilege, output validation, misuse controls, resilience, incident response, and safe shutdown.
Competence, time, context, authority, independence, interface, escalation, override, stop authority, and protection from automation bias.
Ownership, documentation, vendor terms, inventory, change gates, monitoring, audit trail, reassessment, exit, and decommissioning.
Use an evidence ladder
| Status | Meaning | Evidence expectation |
|---|---|---|
| Verified | The control is implemented and supports the defined system and context. | Current applicable document, configuration, test, observation, record, commitment, or independent evidence. |
| Partial | Some implementation or evidence exists, but scope, coverage, testing, ownership, or durability is incomplete. | Record the exact gap, consequence, owner, condition, date, and verification method. |
| Planned | The team intends to implement the control, but it cannot yet reduce the assessed risk. | Treat as an action and gate—not as completed mitigation. |
| Unknown | The control has not been established or the evidence has not been reviewed. | Record the information request and block higher-risk use when the uncertainty is material. |
| Missing | The control is absent, contradicted, or unsuitable. | Redesign, replace, negotiate, limit, or stop according to impact. |
Critical gate: a planned human review does not reduce risk until reviewers have adequate time, information, competence, interface, authority, escalation, fallback, and evidence that they can detect and correct expected failures.
Step 6: Select mitigations and recalculate residual risk
Prefer controls that change the design and reduce the impact pathway at its source. Training and warning text are useful, but they should not carry a risk that could be reduced by narrowing purpose, removing sensitive data, reducing authority, improving the process, or keeping a person responsible for the decision.
Use the control hierarchy
- Avoid: do not use AI for the unsuitable purpose, group, data, decision, or condition.
- Narrow: reduce reach, data, authority, automation, permissions, cases, geography, duration, or affected population.
- Redesign: change workflow, model, interface, data, decision rule, review point, fallback, or vendor.
- Prevent: apply access control, minimization, validation, constraints, separation, representative testing, and approvals.
- Detect and intervene: monitor outcomes, subgroups, errors, misuse, incidents, complaints, overrides, and drift against thresholds.
- Correct and remedy: enable correction, appeal, compensation or other remedy, rollback, incident response, learning, and reassessment.
Write each mitigation as an accountable control
Replace “monitor fairness” with a testable statement: “The Responsible AI lead will review false-negative rates and service outcomes for the approved groups every month. A gap above the agreed threshold triggers investigation within five business days and pauses expansion until the owner records corrective action and residual acceptance.”
Impact pathway → control → owner → deadline → evidence → metric → threshold → response → residual rating → approverRe-rate likelihood and, where the control truly limits consequences, severity. Preserve both inherent and residual ratings. Do not reduce the score because a control is planned, a vendor promises a future feature, or the team expects users to “be careful.”
Step 7: Decide, monitor, and reassess
The assessment should end with a decision, conditions, owners, and a date—not an open list of risks. The decision owner must know which risks are accepted, which must be mitigated before the next gate, and which cannot be accepted.
Proceed to bounded pilot
No critical gate is open; scope is limited; controls and evidence support representative testing with monitoring and rollback.
Proceed with controls
Named gaps are manageable through conditions, owners, deadlines, verification, restricted scope, and explicit pilot limits.
Pause and redesign
A material unknown, impact pathway, weak control, or evidence gap requires redesign or specialist review before proceeding.
Do not deploy yet
Impact exceeds tolerance or required safeguards, remedy, human authority, or evidence cannot be made adequate for the context.
Build monitoring around decisions
- Measure intended benefit and operational cost so the organization can confirm the system remains justified.
- Track serious errors, uncertainty, abstention, corrections, fallback, reviewer burden, and failure by relevant case and group.
- Monitor complaints, challenge, appeal, correction, remedy, accessibility, exclusion, and feedback from affected people.
- Monitor security abuse, prompt injection, sensitive disclosure, unsafe actions, resource use, incidents, outages, and recovery.
- Version models, prompts, retrieval, datasets, tools, vendors, permissions, configurations, policies, tests, and decisions.
- Define thresholds for investigation, scope reduction, rollback, temporary suspension, notification, reassessment, and decommissioning.
Set reassessment triggers
Reassess after material changes to purpose, affected people, authority, scale, geography, language, model, data, retrieval, vendor, subprocessor, integration, permissions, contract, law, operating conditions, performance, complaints, or incidents—and at a risk-based recurring interval.
Worked example: customer refund recommendation assistant
An assistant recommends refund outcomes to trained support agents. It uses order history, approved policy content, and agent notes. The agent sees sources and must approve. It cannot issue refunds, close accounts, decide fraud, or handle vulnerable-customer cases. Refunds above €500 follow the existing specialist process.
Impact pathways
| Impact | Scenario | Inherent rating | Selected controls | Residual decision |
|---|---|---|---|---|
| Economic / service access | Outdated policy retrieval produces an incorrect denial, delaying a legitimate refund. | Serious × Possible = 6 | Versioned source set, source display, policy conflict test, mandatory agent approval, specialist escalation, appeal route | Moderate; pilot with threshold and daily review |
| Fairness | Language quality creates more incomplete recommendations for non-native speakers. | Moderate × Probable = 6 | Multilingual test set, outcome comparison, interpreter route, abstention, no adverse inference from writing style | Moderate; restrict languages until thresholds pass |
| Privacy | Agent notes contain unnecessary sensitive information that appears in prompts or logs. | Serious × Possible = 6 | Field minimization, redaction, access controls, retention, training, vendor terms, log review, incident route | Low-to-moderate after verified controls |
| Security | Untrusted customer text attempts to override instructions and trigger unsafe tool behavior. | Serious × Possible = 6 | Text treated as untrusted data, no action permission, injection testing, output validation, monitoring, kill switch | Low-to-moderate for bounded pilot |
Pilot conditions
- Four-week pilot with 10 trained agents and representative cases; no automatic actions.
- Predefined thresholds for serious recommendation error, unsupported policy claim, agent correction, review time, escalation failure, privacy incident, and language disparity.
- Daily exception review during week one, then weekly; immediate suspension for a defined critical incident.
- Customer appeal and existing non-AI service route remain available.
- Expansion requires a new decision record using pilot evidence—not the initial design assumptions.
This example shows why a score is only one part of the assessment. The useful output is the connection between a plausible impact, a specific control, evidence, an owner, a threshold, and a decision.
Copy-ready AI impact assessment template
Use this record for one defined system and deployment context. Link detailed specialist evidence rather than copying sensitive or lengthy documents into the main record.
AI impact assessment record
Common mistakes and decision red flags
| Mistake or red flag | Why it fails | Better approach |
|---|---|---|
| The assessment begins after purchase or launch approval. | Material design, vendor, data, and contract choices are already difficult to change. | Make AIA a gate in discovery, procurement, pilot, deployment, expansion, and change. |
| The team assesses only the model. | Many harms arise from workflow, interface, data, integration, incentives, human review, or authority. | Map the complete socio-technical system and real operating conditions. |
| Only direct users are considered. | People subject to outputs, bystanders, workers, communities, and non-users may carry the impact. | Map direct, indirect, data, decision, and community stakeholders. |
| Consultation is replaced by internal assumptions. | Power, access, dignity, burden, context, and remedy can be invisible to the project team. | Engage affected people or suitable representatives early and accessibly. |
| Planned controls reduce the score. | A future action cannot currently prevent, detect, or correct harm. | Keep it as a condition and gate until implementation and evidence are verified. |
| Average performance hides subgroup or rare failure. | Serious effects may concentrate in smaller groups or low-frequency high-consequence cases. | Test relevant groups, pathways, edge cases, distribution, and consequence. |
| “Human in the loop” is accepted without evidence. | Reviewers may lack time, information, skill, authority, or ability to resist automation bias. | Design and test human oversight as an operational control. |
| A high overall score overrides a critical gap. | Safety, rights, sensitive data, vulnerable people, or autonomous authority may require a hard gate. | Keep critical decision gates independent from weighted averages. |
| The assessment is filed and forgotten. | Models, data, people, vendors, context, performance, and law change after deployment. | Connect monitoring thresholds, incidents, complaints, change, and time to reassessment. |
Frequently asked questions
Methodology and primary sources
This guide combines practical steps from public AI risk and impact resources. Laws, standards, and official guidance change; verify current versions and applicability for your system, sector, and jurisdiction.
- NIST AI Risk Management Framework — the Govern, Map, Measure, and Manage functions and trustworthy AI characteristics.
- NIST AI RMF Govern Playbook — impact-assessment policies, accountability, participation, risk tolerance, and governance practices.
- NIST AI RMF Map Playbook — context, affected people, impacts, oversight needs, and assessment scales.
- NIST AI RMF Measure Playbook — metrics, evaluation, uncertainty, and limits.
- NIST AI RMF Manage Playbook — prioritization, response, monitoring, intervention, and decommissioning.
- Government of Canada Algorithmic Impact Assessment — a public questionnaire-based approach using risk and mitigation questions to determine impact level.
- ICO AI and Data Protection Risk Toolkit — practical support for risks to individuals' rights and freedoms.
- ICO DPIA guidance — likelihood and severity of potential physical, material, and non-material harm in data-protection risk.
Important: this article is general educational material. It is not legal advice, a compliance determination, an official or independent impact assessment, or authorization to purchase or deploy an AI system.