Thursday, September 10, 2026

GRC INSIGHTS: Volume I | Part VI | AI Incident Management - When AI Goes Wrong: Who Responds, Who Decides, and Who Is Accountable?



Executive Summary

AI Governance is often discussed in terms of policies, risk assessments, controls, monitoring, and executive dashboards.

But the real test of governance comes when an AI system behaves unexpectedly.

An AI model may generate harmful or misleading content. A decision-support system may produce systematically biased outcomes. A third-party AI service may expose confidential information. A model may degrade after deployment. An employee may use an unapproved AI tool and unintentionally disclose sensitive data.

At that moment, the organization faces questions that cannot be answered by a policy document alone:

  • What exactly happened?

  • Is this an AI incident or a conventional technology incident?

  • Who has authority to stop the system?

  • Who owns the business impact?

  • Who decides whether the system can continue operating?

  • Which stakeholders must be informed?

  • What evidence must be preserved?

  • How should the risk register, controls, and governance framework be updated?

This leads to a central principle:

AI Governance is not fully mature until the organization knows how to respond when AI fails.

AI incident management should not become a separate risk universe. It should connect with existing Cybersecurity Incident Response, Privacy Incident Management, Business Continuity, Operational Resilience, Legal, Compliance, Third-Party Risk, and Enterprise GRC processes.

The objective is not simply to restore service.

It is to restore control, accountability, transparency, and trust.

1. AI Incidents Are Not Always Traditional IT Incidents

Traditional IT incident management often focuses on questions such as:

  • Is the system available?

  • Has data been corrupted?

  • Has an account been compromised?

  • Has a service been disrupted?

  • How quickly can functionality be restored?

These questions remain important for AI systems.

However, AI introduces additional dimensions.

An AI system may be available and functioning technically while still producing unacceptable outcomes.

For example:

  • A recruitment model may operate normally while disadvantaging a particular group.

  • A customer-service chatbot may remain online while providing inaccurate regulatory advice.

  • A generative AI assistant may respond quickly while exposing confidential information.

  • A fraud model may continue processing transactions while its accuracy deteriorates because of data drift.

  • An autonomous workflow may complete its assigned task while taking an action that exceeds its approved authority.

In each case, the system may not be “down.”

But governance may still have failed.

Therefore:

An AI incident is not limited to a technical failure. It may also involve a harmful outcome, unacceptable behaviour, control failure, or material deviation from the system’s approved purpose.

2. What Should Qualify as an AI Incident?

Organizations should define AI incidents broadly enough to capture meaningful risk, but precisely enough to support consistent reporting.

A practical definition could be:

An AI incident is an event in which an AI system, its use, or its supporting process causes—or has the potential to cause—material harm, policy violation, security or privacy exposure, unacceptable performance, regulatory non-compliance, or loss of human control.

Potential categories include:

Security incidents

  • Prompt injection

  • Model manipulation

  • Unauthorized access to AI systems

  • Data exfiltration through AI interfaces

  • Abuse of AI-enabled functionality

  • Compromise of an AI service provider

Privacy incidents

  • Disclosure of personal information

  • Inappropriate processing of sensitive data

  • Retention or reuse of data beyond approved purposes

  • Exposure of confidential prompts or conversation history

Safety and performance incidents

  • Dangerous or materially incorrect outputs

  • Model degradation

  • Unacceptable error rates

  • Failure to operate within defined limits

  • Unsafe recommendations or actions

Fairness and responsible AI incidents

  • Discriminatory outcomes

  • Systematic bias

  • Inadequate accessibility

  • Lack of required transparency

  • Inappropriate automated decision-making

Governance and compliance incidents

  • Use of an unapproved AI system

  • Operation outside approved risk appetite

  • Missing human oversight

  • Failure to follow an AI policy

  • Unapproved model or prompt changes

  • Failure to complete required monitoring or reporting

Third-party AI incidents

  • Vendor model failure

  • Unexpected service changes

  • Loss of contractual safeguards

  • Third-party data leakage

  • Uncommunicated changes to model behaviour

This classification should align with the organization’s existing incident taxonomy wherever possible.

The goal is integration, not the creation of another disconnected process.

3. The AI Incident Lifecycle

A practical AI incident-management lifecycle can be structured into seven stages:

1. Detect

Identify an unusual event, harmful output, control failure, or emerging concern.

2. Triage

Determine severity, scope, affected stakeholders, and whether immediate containment is required.

3. Contain

Prevent further harm by restricting, suspending, isolating, or disabling the relevant AI capability.

4. Investigate

Establish what happened, why it happened, and whether the issue originated in the model, data, prompt, workflow, user, vendor, or control environment.

5. Remediate

Address the underlying cause and restore the system to an acceptable operating state.

6. Recover

Resume operations only when the appropriate technical, business, risk, legal, and governance conditions have been satisfied.

7. Learn

Update risk assessments, controls, monitoring, training, documentation, and governance decisions based on the incident.

The final stage is essential.

An incident should not be considered fully closed merely because the system is operational again.

Operational recovery is not the same as governance recovery.

4. Who Owns the Incident?

One of the most common weaknesses in incident management is unclear ownership.

AI incidents often cross organizational boundaries.

A single event may involve:

  • The business owner

  • The AI product owner

  • Data Governance

  • Cybersecurity

  • Privacy

  • Legal

  • Compliance

  • Model Risk

  • Enterprise Risk

  • Procurement

  • The third-party provider

  • Communications

  • Senior management

If ownership is unclear, response becomes slow and fragmented.

A useful distinction is between four different responsibilities.

Incident Coordinator

Coordinates the response, maintains the incident record, and ensures that actions are tracked.

Technical Owner

Investigates the system, model, data, infrastructure, or integration involved.

Business Owner

Assesses the business impact and decides what operational outcomes are acceptable.

Accountable Executive

Owns the material risk decision, including whether the AI system should remain operational, be restricted, or be withdrawn.

These roles may be performed by different individuals depending on the organization.

The important point is that they should be explicitly defined before an incident occurs.

NIST's AI RMF emphasizes documented roles, responsibilities, and lines of communication for managing AI risks. It also calls for processes covering incident response, recovery, change management, and communication.

5. The First Decision: Continue, Restrict, or Stop?

During an AI incident, the organization may need to make a rapid operational decision.

A practical decision framework is:

Continue

Use this when:

  • The issue is low severity.

  • The impact is understood.

  • Existing controls remain effective.

  • Monitoring is sufficient.

  • The risk remains within approved tolerance.

Restrict

Use this when:

  • The system can operate safely only within narrower limits.

  • Certain users, data types, geographies, or decisions must be excluded.

  • Human review needs to be strengthened.

  • Particular functions need to be disabled.

Suspend

Use this when:

  • Material harm may continue.

  • The root cause is not yet understood.

  • Critical controls have failed.

  • The system is operating outside approved risk appetite.

Decommission or Withdraw

Use this when:

  • The system cannot be brought within acceptable risk.

  • The intended use is no longer appropriate.

  • The risk cannot be effectively mitigated.

  • The system has become unsuitable for the business or regulatory context.

This decision should not be left solely to a technical team.

A technical team can explain whether a system is functioning.

It may not have the authority to decide whether the organization should accept the associated business, legal, ethical, or regulatory risk.

6. AI Incident Severity Should Reflect Impact, Not Just Downtime

Traditional technology severity models often prioritize availability and infrastructure impact.

AI incident severity should consider a broader set of dimensions.

A practical AI incident severity assessment could include:

Impact on people

Could the incident affect individuals’ rights, safety, access to services, employment, finances, or reputation?

Impact on data

Was confidential, personal, regulated, or proprietary information exposed?

Impact on business operations

Could the incident disrupt critical processes or cause financial loss?

Impact on trust

Could the incident damage customer, employee, partner, or public confidence?

Regulatory impact

Could the incident trigger notification, reporting, investigation, or enforcement obligations?

Scale and duration

How many people, transactions, systems, or decisions may be affected?

Reversibility

Can the impact be corrected, or are the consequences permanent?

Degree of human control

Was a meaningful human review available, or did the system act autonomously?

An incident involving a chatbot outage may be inconvenient.

An incident involving an AI system that makes unsafe or discriminatory decisions may be materially more serious even if the system remains fully available.

7. Containment Must Include More Than Technical Controls

Technical containment may include:

  • Disabling an API

  • Blocking a model

  • Restricting access

  • Rolling back a deployment

  • Reverting a prompt or configuration

  • Isolating affected data

  • Suspending automated actions

  • Requiring human approval

But AI incidents may also require non-technical containment.

Examples include:

  • Pausing affected business decisions

  • Notifying impacted customers or employees

  • Rechecking decisions made by the system

  • Suspending a vendor relationship

  • Preserving evidence

  • Engaging Legal or Privacy teams

  • Updating frontline staff

  • Introducing temporary manual processes

  • Reviewing previously generated outputs

This is especially important where AI has influenced decisions that cannot simply be “rolled back.”

For example, if an AI system incorrectly rejected applications, flagged customers, or generated inaccurate advice, the organization may need to identify and review the affected population.

The question is not only:

“How do we stop the system?”

It is also:

“How do we address the consequences of what the system has already done?”

8. Investigation Requires an AI-Specific Evidence Trail

An AI incident investigation may require evidence that is not routinely captured in conventional systems.

Depending on the use case, relevant evidence may include:

  • Model version

  • Prompt or instruction version

  • Input data

  • Output generated

  • Retrieval sources

  • System configuration

  • Model parameters

  • User identity and access context

  • Timestamp

  • Human review or override records

  • Tool calls initiated by the AI system

  • External system actions

  • Monitoring alerts

  • Model-performance data

  • Data and model changes

  • Vendor notifications

  • Relevant policy and approval records

Without this evidence, the organization may be unable to answer basic questions:

  • What did the AI receive?

  • What did it produce?

  • What did the user see?

  • What action did the system take?

  • Which version of the model was operating?

  • Was human oversight applied?

  • Was the system operating within its approved purpose?

This creates a direct connection between AI Governance and AI observability.

If the organization cannot reconstruct what happened, it may not be able to demonstrate that it governed the system effectively.

9. Root Cause Analysis Must Look Beyond the Model

A common mistake is to assume that every AI incident is caused by the model.

In reality, the root cause may exist elsewhere.

Potential causes include:

Data

  • Poor-quality training data

  • Incomplete data

  • Data drift

  • Inappropriate data sources

  • Incorrect data labelling

Model

  • Model limitations

  • Unexpected behaviour

  • Inadequate testing

  • Performance degradation

  • Poor generalization

Prompt or configuration

  • Unsafe instructions

  • Inadequate system prompts

  • Incorrect thresholds

  • Misconfigured guardrails

Workflow

  • AI used for an unsuitable purpose

  • Missing human review

  • Excessive automation

  • Poor exception handling

People

  • Inadequate training

  • Misuse

  • Misinterpretation of outputs

  • Failure to follow procedures

Technology

  • Integration defects

  • Access-control failures

  • Logging gaps

  • Infrastructure issues

Third party

  • Uncommunicated model changes

  • Vendor outage

  • Contractual control failure

  • Inadequate supplier transparency

Therefore, AI incident root cause analysis should examine the complete socio-technical system, not only the model.

10. Communication Is Part of Incident Management

AI incidents can create confusion because stakeholders may not understand what the system did, why it did it, or whether the organization remains in control.

Communication should therefore be planned in advance.

Relevant stakeholders may include:

  • Senior management

  • The Board or relevant Board committee

  • Employees

  • Customers

  • Regulators

  • Business partners

  • Affected individuals

  • Vendors

  • Internal audit

  • Legal and Compliance

  • Public relations teams

The communication should be:

Accurate

Avoid speculation and unsupported explanations.

Timely

Do not delay necessary escalation while waiting for a perfect root-cause analysis.

Proportionate

Match the communication to the severity and impact.

Transparent

Explain what is known, what is not known, and what action is being taken.

Accountable

Identify the responsible function and the next steps.

NIST's AI RMF includes communicating incidents and errors to relevant AI actors, including affected communities where appropriate.

For high-risk AI systems covered by the EU AI Act, the regulation also includes post-market monitoring and serious-incident reporting obligations for providers. The exact obligations depend on the system, role, jurisdiction, and applicable legal requirements.

11. Recovery Requires Explicit Exit Criteria

A system should not automatically return to production simply because the immediate issue appears to be resolved.

Recovery criteria should be defined according to the incident’s severity and nature.

Possible criteria include:

  • The immediate threat has been contained.

  • The root cause is sufficiently understood.

  • Required controls have been restored.

  • Performance is within approved thresholds.

  • Human oversight is functioning.

  • Affected data has been secured.

  • Required testing has been completed.

  • Residual risk has been reassessed.

  • The appropriate risk owner has accepted the residual risk.

  • Legal, Privacy, Compliance, or regulatory obligations have been addressed.

  • Business and executive approval has been obtained where required.

For high-impact AI systems, recovery may also require:

  • Revalidation of the model

  • Independent review

  • Reassessment of fairness or safety

  • Rechecking affected decisions

  • Additional monitoring

  • Temporary operating restrictions

The recovery decision should be documented.

Otherwise, the organization may restore the system technically without demonstrating that it has restored acceptable governance conditions.

12. Every Incident Should Update the Risk Register

An AI incident should not disappear into an incident-management platform after closure.

It should feed back into the broader governance system.

The organization should consider whether the incident requires updates to:

  • The AI Risk Register

  • The AI system inventory

  • The risk classification

  • The control library

  • The monitoring plan

  • The AI impact assessment

  • The vendor assessment

  • The risk appetite decision

  • The business continuity plan

  • The AI policy

  • Employee training

  • Contractual requirements

  • Model documentation

  • Testing procedures

This is how incident management becomes a learning mechanism.

A mature organization does not simply ask:

“How do we prevent this exact incident from happening again?”

It also asks:

“What does this incident reveal about weaknesses in our governance system?”

13. The AI Incident Learning Loop

The relationship between AI Governance and incident management can be represented as a continuous loop:

Governance requirements

AI system design and deployment

Monitoring and detection

Incident response

Root cause analysis

Risk and control updates

Improved governance requirements

More resilient AI systems

This is important because AI risks are not static.

NIST's recent work on monitoring deployed AI systems highlights the need to monitor functionality, operations, human factors, security, compliance, and broader impacts after deployment. It also identifies challenges such as detecting performance degradation, fragmented logging, emerging risks, and balancing automated monitoring with human-validated monitoring.

The implication is clear:

Incident management should not be the end of the governance process. It should improve the next version of the governance process.

14. Metrics That Matter in AI Incident Management

Building on the previous GRC Insights article about AI Governance metrics, organizations should monitor incident-management effectiveness.

Useful indicators include:

Detection

  • Number of AI incidents detected internally

  • Percentage detected through monitoring

  • Average time to detect

  • Percentage of incidents reported by employees or users

Response

  • Average time to triage

  • Average time to contain

  • Percentage of incidents escalated within required timelines

  • Percentage of incidents with an assigned accountable owner

Investigation

  • Percentage with complete evidence records

  • Average time to establish root cause

  • Percentage requiring third-party investigation

  • Percentage involving human oversight failure

Recovery

  • Average time to restore acceptable operation

  • Percentage reopened after closure

  • Percentage requiring review of previous decisions

  • Percentage with documented residual-risk acceptance

Learning

  • Percentage resulting in control improvements

  • Percentage resulting in updated risk assessments

  • Number of recurring incidents

  • Number of new risks identified after incidents

  • Percentage of lessons learned implemented

However, incident volume should not be interpreted in isolation.

A rise in reported incidents may indicate worsening AI performance.

It may also indicate better detection, stronger reporting culture, or improved transparency.

This is why:

A mature incident-management program measures both the incidents and the organization’s ability to detect, respond, learn, and improve.

15. The AI Incident Readiness Test

Before deploying a material AI system, leadership should be able to answer the following questions:

  1. What constitutes an incident for this system?

  2. Who can report an incident?

  3. Who owns the response?

  4. Who has authority to suspend the system?

  5. What thresholds trigger escalation?

  6. What evidence will be captured?

  7. How will affected decisions or outputs be identified?

  8. What stakeholders must be informed?

  9. What are the recovery criteria?

  10. How will the incident update the risk register and control environment?

If the organization cannot answer these questions, the system may not be ready for deployment.

16. The Three Accountability Questions

During every material AI incident, leadership should ask three questions.

1. Who is accountable for the system?

This is the business or executive accountability question.

2. Who is accountable for the response?

This is the incident-management question.

3. Who is accountable for the decision to resume operation?

This is the risk-acceptance and recovery question.

These responsibilities may sit with different people.

That is acceptable.

What is not acceptable is allowing them to remain undefined.

17. AI Incident Management Should Extend Enterprise GRC

AI incident management should connect with existing organizational capabilities.

AI Incident Area

Existing Enterprise Capability

Model or system failure

IT Service Management

Cyberattack or prompt injection

Cybersecurity Incident Response

Personal-data exposure

Privacy Incident Management

Regulatory concern

Legal and Compliance

Third-party model failure

Third-Party Risk Management

Business disruption

Business Continuity and Resilience

Material risk acceptance

Enterprise Risk Management

Control failure

Internal Control and Assurance

Customer impact

Customer Complaints and Conduct Risk

Repeated failure

Internal Audit and Continual Improvement

This is the key governance principle:

AI incident management should strengthen Enterprise GRC—not create a parallel incident universe.

18. What the Board Should Ask

A Board or executive committee does not need every technical detail.

But it should ask whether the organization is prepared to respond effectively.

Useful questions include:

  • What types of AI incidents are most material to our business?

  • Which AI systems could create significant harm if they failed?

  • Who has authority to suspend those systems?

  • Have we tested our AI incident-response process?

  • Can we reconstruct what an AI system did?

  • How quickly can we identify affected customers, employees, or decisions?

  • Are AI incidents integrated with existing enterprise incident processes?

  • How do we know whether our monitoring is effective?

  • Have previous incidents resulted in measurable improvements?

  • Are we prepared to communicate with regulators or affected stakeholders where required?

The Board should not only ask whether an incident-response plan exists.

It should ask:

“Have we demonstrated that the plan works?”

Final Thoughts

AI Governance is often strongest on paper before an AI system is deployed.

The real test comes later.

When the system produces an unexpected result.

When a control fails.

When a vendor changes a model.

When confidential data is exposed.

When a decision is challenged.

When the organization must decide whether to continue, restrict, or stop the system.

At that moment, governance becomes operational.

The organization must know:

  • What happened

  • Who is responsible

  • Who has authority to act

  • What evidence must be preserved

  • Who needs to be informed

  • What risk remains

  • When the system can safely resume

  • What must change to prevent recurrence

The most important lesson is this:

An AI incident is not only a failure of technology. It is a test of the organization’s ability to exercise control, accountability, and judgment.

A mature AI Governance capability therefore needs more than policies, risk registers, maturity models, and dashboards.

It needs a tested ability to respond when reality does not behave as expected.

Because the ultimate measure of responsible AI is not whether the organization can claim that its systems are trustworthy.

It is whether the organization can detect, explain, contain, remediate, and learn from the moments when trust is put at risk.

Looking Ahead

The GRC Insights series has now progressed through six connected themes:

Part I — People Why Every Employee Is an AI Data Steward

Part II — Risk The AI Governance Blind Spot

Part III — Operationalization Building an AI Risk Register

Part IV — Maturity The AI Governance Maturity Model

Part V — Measurement AI Governance Metrics That Matter

Part VI — Incident Management When AI Goes Wrong: Who Responds, Who Decides, and Who Is Accountable?

The next question is a natural one:

How can organizations demonstrate that their AI Governance framework is not only designed well, but operating effectively?

That brings us to the next article:

AI Governance Assurance

From Policy to Proof: How Organizations Can Demonstrate Responsible AI Governance

The next edition will explore assurance, internal audit, control testing, evidence, independent review, and how organizations can build confidence that their AI Governance framework is operating as intended.

Sources & Further Reading

The principles discussed in this article were informed by established AI risk-management frameworks, standards, and regulatory guidance.

The AI Incident Lifecycle, Three Accountability Questions, and AI Incident Readiness Test are practical models developed for the GRC Insights series. They are not official frameworks issued by NIST, ISO, or any regulator.

1. NIST — Artificial Intelligence Risk Management Framework

The NIST AI Risk Management Framework provides a voluntary structure for managing risks associated with AI systems across their lifecycle.

Its Manage function includes risk treatment, response and recovery planning, post-deployment monitoring, incident response, communication, and continual improvement.

Official source: NIST AI Risk Management Framework 

2. NIST AI RMF Core and Playbook

The NIST AI RMF identifies the importance of clear accountability structures, documented responsibilities, monitoring, incident response, recovery, change management, and communication with relevant AI actors.

These principles informed the article’s emphasis on clearly defined ownership and the integration of AI incident management with Enterprise GRC.

Official source: NIST AI RMF Core 

3. NIST — Challenges to the Monitoring of Deployed AI Systems

NIST’s 2026 report examines the challenges of monitoring AI systems after deployment, including functionality, operations, human factors, security, compliance, and broader impacts.

The report highlights issues such as performance degradation, fragmented logging, emerging risks, incident-sharing gaps, and the need to balance automated monitoring with human-validated monitoring.

Official source: NIST — Challenges to the Monitoring of Deployed AI Systems 

4. ISO/IEC 42001:2023 — Artificial Intelligence Management System

ISO/IEC 42001 specifies requirements for establishing, implementing, maintaining, and continually improving an Artificial Intelligence Management System.

Its management-system approach supports the article’s emphasis on documented responsibilities, performance evaluation, continual improvement, and evidence-based governance.

Official source: ISO/IEC 42001:2023 

5. European Union — EU AI Act

The EU AI Act establishes a risk-based regulatory framework for AI.

Its provisions concerning post-market monitoring and serious-incident reporting reinforce the importance of monitoring AI systems throughout their lifecycle and responding appropriately when serious incidents occur. The precise obligations depend on the AI system, the organization’s role, and the applicable legal requirements.

Official source: European Union — Artificial Intelligence Act 

6. NIST AI RMF to ISO/IEC 42001 Crosswalk

The NIST-to-ISO/IEC 42001 crosswalk demonstrates how NIST AI RMF practices relating to monitoring, incident response, recovery, communication, and continual improvement align with provisions of ISO/IEC 42001.

Official source: NIST AI RMF to ISO/IEC 42001 Crosswalk 

GRC INSIGHTS: Volume I | Part VI | AI Incident Management - When AI Goes Wrong: Who Responds, Who Decides, and Who Is Accountable?

Executive Summary AI Governance is often discussed in terms of policies, risk assessments, controls, monitoring, and executive dashboards. B...