Skip to main content

AIGO — AI Governance Operating Framework

AI Incident Management Procedure

Version: 0.1
Status: Draft
Working Name: AIGO
Full Name: AI Governance Operating Framework
Document Identifier: AIGO-PROC-008

1. Purpose

This procedure defines the process for identifying, reporting, assessing, containing, investigating, resolving, documenting, and learning from incidents involving AI systems. The procedure establishes a consistent governance process for managing actual or suspected events that may adversely affect the confidentiality, integrity, availability, safety, privacy, security, compliance, performance, or intended operation of an AI system.

2. Scope

This procedure applies to incidents involving AI systems within the organization’s AIGO governance scope. It may apply to:
  • AI models;
  • AI applications;
  • generative AI systems;
  • AI agents;
  • AI-enabled business processes;
  • third-party AI services;
  • AI data pipelines;
  • AI infrastructure;
  • AI governance controls; and
  • supporting systems whose failure materially affects an AI system.

3. Objectives

The objectives of AI incident management are to:
  • identify incidents promptly;
  • assess potential impact;
  • establish appropriate severity;
  • contain harmful effects;
  • protect affected people and systems;
  • restore controlled operation;
  • meet reporting obligations;
  • preserve evidence;
  • identify root causes;
  • implement corrective actions; and
  • prevent recurrence.

4. Incident Management Principles

AI incidents should be managed according to the following principles:
  • prompt identification;
  • proportionate response;
  • risk-based escalation;
  • protection of affected individuals;
  • evidence preservation;
  • clear accountability;
  • documented decision-making;
  • appropriate communication; and
  • continuous improvement.

5. Incident Definition

An AI incident is an event or condition involving an AI system that results in, or may reasonably result in:
  • harm;
  • material risk;
  • security compromise;
  • privacy violation;
  • unsafe behavior;
  • significant service disruption;
  • material control failure;
  • regulatory non-compliance; or
  • significant deviation from intended operation.

6. Incident and Event Distinction

Not every operational event is an incident. An event may become an incident when its actual or potential impact exceeds defined organizational thresholds. Organizations should establish criteria for distinguishing:
  • normal events;
  • operational issues;
  • control exceptions;
  • incidents;
  • major incidents; and
  • crises.

7. Incident Sources

AI incidents may be identified through:
  • users;
  • employees;
  • customers;
  • monitoring systems;
  • security monitoring;
  • privacy monitoring;
  • control testing;
  • audits;
  • assurance activities;
  • suppliers;
  • regulators; and
  • automated detection.

8. Incident Triggers

Incident management should be initiated when there is evidence or reasonable suspicion of a material adverse event. Triggers may include:
  • unsafe AI output;
  • harmful AI action;
  • unauthorized access;
  • data leakage;
  • privacy violation;
  • security compromise;
  • material model failure;
  • control failure;
  • significant service disruption;
  • regulatory concern; or
  • unexpected high-impact behavior.

9. Incident Reporting

AI incidents should be reported through defined organizational channels. Reports should contain, where known:
  • reporter;
  • date and time;
  • AI system;
  • description;
  • observed behavior;
  • affected users;
  • potential impact;
  • available evidence; and
  • immediate actions taken.

10. Initial Triage

Reported incidents should undergo initial triage. Triage should determine:
  • whether an incident exists;
  • affected AI system;
  • potential severity;
  • immediate risks;
  • required containment;
  • required escalation; and
  • responsible incident owner.

11. Incident Severity

The organization should establish an incident severity model. A representative model may include:
  1. Low.
  2. Moderate.
  3. High.
  4. Critical.
Severity should be determined using defined criteria.

12. Severity Factors

Incident severity may consider:
  • actual harm;
  • potential harm;
  • number of affected individuals;
  • sensitivity of data;
  • system criticality;
  • duration;
  • geographic scope;
  • regulatory impact;
  • financial impact;
  • reputational impact; and
  • reversibility.

13. Critical AI Incidents

Critical incidents may include events involving:
  • serious or potentially serious harm;
  • significant safety risk;
  • major privacy breach;
  • major security compromise;
  • widespread system failure;
  • material regulatory exposure; or
  • uncontrolled autonomous behavior.
Critical incidents should receive immediate escalation.

14. Incident Owner

Each material incident should have an assigned incident owner. The incident owner is responsible for:
  • coordinating response;
  • maintaining the incident record;
  • coordinating stakeholders;
  • tracking decisions;
  • ensuring escalation; and
  • confirming closure.

15. Incident Response Team

Depending on severity, the response may involve:
  • AI governance;
  • AI system owner;
  • security;
  • privacy;
  • legal;
  • compliance;
  • risk management;
  • technical operations;
  • communications;
  • business leadership; and
  • external specialists.

16. Immediate Containment

Where necessary, immediate containment should be performed to limit impact. Containment actions may include:
  • disabling functionality;
  • restricting access;
  • stopping automated actions;
  • isolating systems;
  • disabling integrations;
  • reverting to a previous model;
  • increasing human oversight; or
  • suspending the AI system.

17. Safety Protection

Where an AI incident may cause physical or significant individual harm, protection of affected people should take priority. Actions may include:
  • stopping affected operations;
  • increasing human intervention;
  • restricting system capabilities;
  • initiating emergency procedures; and
  • notifying appropriate authorities or responsible personnel.

18. Security Containment

Security-related AI incidents should be coordinated with applicable security incident processes. Actions may include:
  • credential revocation;
  • access restriction;
  • isolation;
  • forensic preservation;
  • malicious activity blocking; and
  • security monitoring escalation.

19. Privacy Containment

Privacy-related incidents should be coordinated with applicable privacy incident processes. Actions may include:
  • restricting data access;
  • stopping processing;
  • securing exposed information;
  • identifying affected data;
  • preserving evidence; and
  • initiating required notification assessment.

20. Incident Investigation

Material incidents should be investigated to determine:
  • what occurred;
  • when it occurred;
  • how it occurred;
  • systems involved;
  • data involved;
  • people affected;
  • controls involved;
  • contributing factors; and
  • potential root causes.

21. Evidence Preservation

Evidence should be preserved according to applicable requirements. Evidence may include:
  • logs;
  • model versions;
  • prompts;
  • outputs;
  • configurations;
  • access records;
  • system events;
  • monitoring data;
  • communications; and
  • relevant documentation.

22. AI-Specific Evidence

Where relevant, incident investigation should preserve AI-specific information such as:
  • model identifier;
  • model version;
  • system prompt;
  • user prompt;
  • generated output;
  • tool calls;
  • agent actions;
  • model configuration;
  • retrieved data;
  • relevant context;
  • safety controls; and
  • human interventions.

23. Root Cause Analysis

Material incidents should undergo root cause analysis where appropriate. Potential root causes may include:
  • model behavior;
  • data quality;
  • configuration;
  • software defect;
  • control failure;
  • human error;
  • inadequate monitoring;
  • third-party failure;
  • process weakness; or
  • governance failure.

24. Contributing Factors

The investigation should distinguish between:
  • root cause;
  • contributing factors;
  • triggering event; and
  • underlying governance weaknesses.
This distinction should support effective remediation.

25. Control Failure Analysis

Where controls failed or were bypassed, the investigation should determine:
  • which control failed;
  • whether the control was designed appropriately;
  • whether it was implemented;
  • whether it operated correctly;
  • whether evidence existed; and
  • whether compensating controls were effective.

26. Risk Reassessment

An incident should trigger risk reassessment where it indicates a material change in the AI system’s risk profile. Risk reassessment should consider:
  • likelihood;
  • impact;
  • affected populations;
  • existing controls;
  • residual risk; and
  • newly identified risks.

27. Classification Reassessment

The AI system’s governance classification should be reviewed where the incident indicates that its existing classification may no longer be appropriate.

28. Containment Decision

The incident owner should determine whether the AI system should:
  • continue operating;
  • operate with restrictions;
  • operate under enhanced monitoring;
  • revert to a previous version;
  • be temporarily suspended; or
  • be permanently withdrawn.

29. Business Continuity

Where an AI system is operationally critical, incident response should consider continuity requirements. Alternative measures may include:
  • manual processes;
  • fallback systems;
  • alternative providers;
  • reduced functionality; and
  • temporary suspension of affected services.

30. Communication

Incident communication should be proportionate to severity. Potential stakeholders include:
  • affected users;
  • business owners;
  • executives;
  • security;
  • privacy;
  • legal;
  • regulators;
  • customers;
  • suppliers; and
  • other affected parties.

31. Regulatory Notification

Where required, the organization should assess whether an incident triggers notification or reporting obligations. The assessment should be performed by appropriately authorized personnel.

32. External Communication

External communications should be coordinated through authorized functions. Communications should be:
  • accurate;
  • timely;
  • proportionate;
  • consistent; and
  • appropriately documented.

33. Incident Remediation

Corrective actions should address identified weaknesses. Remediation may include:
  • model changes;
  • control improvements;
  • configuration changes;
  • additional testing;
  • monitoring improvements;
  • training;
  • process changes; and
  • governance changes.

34. Corrective Action Plan

Material corrective actions should identify:
  • finding;
  • required action;
  • owner;
  • priority;
  • target date;
  • evidence;
  • verification; and
  • closure criteria.

35. Preventive Actions

Where appropriate, preventive actions should be implemented to reduce recurrence. Preventive actions may address:
  • system design;
  • controls;
  • testing;
  • monitoring;
  • training;
  • procedures;
  • governance; and
  • supplier management.

36. Incident Closure

An incident should only be closed when:
  • containment is complete;
  • required investigation is complete;
  • immediate risks are addressed;
  • required notifications are completed;
  • corrective actions are assigned;
  • residual risk is understood; and
  • closure is authorized.

37. Incident Review

Material incidents should receive a formal post-incident review. The review should consider:
  • response effectiveness;
  • containment;
  • communication;
  • decision-making;
  • controls;
  • monitoring;
  • root cause;
  • remediation; and
  • lessons learned.

38. Lessons Learned

Lessons learned should be incorporated into the AIGO governance system. Potential updates include:
  • controls;
  • risk assessments;
  • system profiles;
  • procedures;
  • monitoring;
  • training;
  • architecture;
  • testing; and
  • governance requirements.

39. Incident Records

Incident records should be maintained according to applicable retention requirements. Records should include:
  • incident identifier;
  • date and time;
  • AI system;
  • severity;
  • description;
  • impact;
  • containment;
  • investigation;
  • evidence;
  • root cause;
  • actions;
  • notifications;
  • approvals; and
  • closure.

40. Incident Traceability

Incident records should maintain traceability to relevant governance artifacts. The organization should be able to demonstrate: Incident → AI System → Risk → Controls → Response → Investigation → Remediation → Lessons Learned

41. Incident Monitoring

Open incidents should be actively monitored until closure. Monitoring should track:
  • actions;
  • owners;
  • deadlines;
  • severity;
  • residual risk;
  • outstanding decisions; and
  • escalation requirements.

42. Incident Escalation

Incidents should be escalated when:
  • severity increases;
  • containment fails;
  • new affected parties are identified;
  • regulatory exposure increases;
  • risk exceeds tolerance;
  • required actions are delayed; or
  • the incident becomes a major or critical event.

43. Incident Metrics

Organizations may establish incident management metrics. Examples include:
  • incident count;
  • incident severity;
  • time to detection;
  • time to containment;
  • time to resolution;
  • repeat incidents;
  • control-related incidents;
  • unauthorized AI actions; and
  • overdue corrective actions.

Incident trends should be periodically reviewed to identify recurring weaknesses. Trend analysis may identify:
  • recurring model failures;
  • repeated control failures;
  • repeated security issues;
  • recurring data problems;
  • supplier issues; and
  • systemic governance weaknesses.

45. Third-Party Incidents

Third-party AI incidents should be managed according to contractual and organizational requirements. The organization should determine:
  • provider notification;
  • internal escalation;
  • customer impact;
  • regulatory implications;
  • containment;
  • provider remediation; and
  • ongoing monitoring.

46. Incident and Change Management

Material incident remediation may require formal change management. Changes resulting from an incident should be assessed according to the AIGO change management procedure.

47. Incident and Risk Management

Incident findings should feed into AI risk management. Where an incident demonstrates that an existing risk assessment is inaccurate or incomplete, the relevant risk record should be updated.

48. Incident and Control Management

Incident findings should feed into control assessment. Controls that failed, were insufficient, or were bypassed should be reviewed and improved.

49. Incident and Monitoring

Incident findings should be used to improve monitoring. Where appropriate, new:
  • indicators;
  • thresholds;
  • alerts;
  • detection mechanisms; and
  • review requirements
should be established.

50. Responsibilities

Incident Owner
  • coordinate incident response;
  • maintain the incident record;
  • coordinate investigation;
  • track actions;
  • escalate appropriately; and
  • confirm closure.
AI System Owner
  • provide system information;
  • support containment;
  • coordinate technical remediation;
  • reassess system risk; and
  • implement required corrective actions.
AI Governance Function
  • oversee governance implications;
  • coordinate escalation;
  • maintain governance records;
  • monitor material incidents; and
  • report significant incidents.
Specialist Functions
  • provide security, privacy, legal, compliance, risk, technical, safety, communications, or other specialist support as required.

51. Incident Management Workflow

The standard workflow should be:
  1. Detect or receive incident report.
  2. Register incident.
  3. Perform initial triage.
  4. Determine severity.
  5. Assign incident owner.
  6. Initiate containment.
  7. Escalate where required.
  8. Preserve evidence.
  9. Investigate incident.
  10. Determine root cause and contributing factors.
  11. Assess risk and control implications.
  12. Determine notification requirements.
  13. Implement remediation.
  14. Validate remediation.
  15. Conduct post-incident review.
  16. Update governance artifacts.
  17. Record lessons learned.
  18. Obtain closure approval.
  19. Close incident.

52. Continuous Improvement

The incident management process should be improved based on:
  • incidents;
  • near misses;
  • response experience;
  • assurance findings;
  • audit findings;
  • monitoring results;
  • stakeholder feedback;
  • regulatory developments; and
  • changes in AI technology.

53. Procedure Review

This procedure should be reviewed periodically and when material changes occur. Review triggers may include:
  • significant incidents;
  • changes to AIGO requirements;
  • regulatory developments;
  • assurance findings;
  • changes to organizational incident management;
  • changes in AI technology; and
  • implementation experience.
Material changes should be versioned and approved according to applicable document governance requirements.

54. Procedure Status

Document: AIGO AI Incident Management Procedure Version: 0.1 Status: Draft Working Name: AIGO Full Name: AI Governance Operating Framework Document Identifier: AIGO-PROC-008 Document Type: Operational Procedure This procedure establishes the operational process for identifying, managing, investigating, resolving, and learning from AI incidents throughout the AI governance lifecycle.