AIGO — AI Governance Operating Framework
AI Incident Management Procedure
Version: 0.1Status: Draft
Working Name: AIGO
Full Name: AI Governance Operating Framework
Document Identifier: AIGO-PROC-008
1. Purpose
This procedure defines the process for identifying, reporting, assessing, containing, investigating, resolving, documenting, and learning from incidents involving AI systems. The procedure establishes a consistent governance process for managing actual or suspected events that may adversely affect the confidentiality, integrity, availability, safety, privacy, security, compliance, performance, or intended operation of an AI system.2. Scope
This procedure applies to incidents involving AI systems within the organization’s AIGO governance scope. It may apply to:- AI models;
- AI applications;
- generative AI systems;
- AI agents;
- AI-enabled business processes;
- third-party AI services;
- AI data pipelines;
- AI infrastructure;
- AI governance controls; and
- supporting systems whose failure materially affects an AI system.
3. Objectives
The objectives of AI incident management are to:- identify incidents promptly;
- assess potential impact;
- establish appropriate severity;
- contain harmful effects;
- protect affected people and systems;
- restore controlled operation;
- meet reporting obligations;
- preserve evidence;
- identify root causes;
- implement corrective actions; and
- prevent recurrence.
4. Incident Management Principles
AI incidents should be managed according to the following principles:- prompt identification;
- proportionate response;
- risk-based escalation;
- protection of affected individuals;
- evidence preservation;
- clear accountability;
- documented decision-making;
- appropriate communication; and
- continuous improvement.
5. Incident Definition
An AI incident is an event or condition involving an AI system that results in, or may reasonably result in:- harm;
- material risk;
- security compromise;
- privacy violation;
- unsafe behavior;
- significant service disruption;
- material control failure;
- regulatory non-compliance; or
- significant deviation from intended operation.
6. Incident and Event Distinction
Not every operational event is an incident. An event may become an incident when its actual or potential impact exceeds defined organizational thresholds. Organizations should establish criteria for distinguishing:- normal events;
- operational issues;
- control exceptions;
- incidents;
- major incidents; and
- crises.
7. Incident Sources
AI incidents may be identified through:- users;
- employees;
- customers;
- monitoring systems;
- security monitoring;
- privacy monitoring;
- control testing;
- audits;
- assurance activities;
- suppliers;
- regulators; and
- automated detection.
8. Incident Triggers
Incident management should be initiated when there is evidence or reasonable suspicion of a material adverse event. Triggers may include:- unsafe AI output;
- harmful AI action;
- unauthorized access;
- data leakage;
- privacy violation;
- security compromise;
- material model failure;
- control failure;
- significant service disruption;
- regulatory concern; or
- unexpected high-impact behavior.
9. Incident Reporting
AI incidents should be reported through defined organizational channels. Reports should contain, where known:- reporter;
- date and time;
- AI system;
- description;
- observed behavior;
- affected users;
- potential impact;
- available evidence; and
- immediate actions taken.
10. Initial Triage
Reported incidents should undergo initial triage. Triage should determine:- whether an incident exists;
- affected AI system;
- potential severity;
- immediate risks;
- required containment;
- required escalation; and
- responsible incident owner.
11. Incident Severity
The organization should establish an incident severity model. A representative model may include:- Low.
- Moderate.
- High.
- Critical.
12. Severity Factors
Incident severity may consider:- actual harm;
- potential harm;
- number of affected individuals;
- sensitivity of data;
- system criticality;
- duration;
- geographic scope;
- regulatory impact;
- financial impact;
- reputational impact; and
- reversibility.
13. Critical AI Incidents
Critical incidents may include events involving:- serious or potentially serious harm;
- significant safety risk;
- major privacy breach;
- major security compromise;
- widespread system failure;
- material regulatory exposure; or
- uncontrolled autonomous behavior.
14. Incident Owner
Each material incident should have an assigned incident owner. The incident owner is responsible for:- coordinating response;
- maintaining the incident record;
- coordinating stakeholders;
- tracking decisions;
- ensuring escalation; and
- confirming closure.
15. Incident Response Team
Depending on severity, the response may involve:- AI governance;
- AI system owner;
- security;
- privacy;
- legal;
- compliance;
- risk management;
- technical operations;
- communications;
- business leadership; and
- external specialists.
16. Immediate Containment
Where necessary, immediate containment should be performed to limit impact. Containment actions may include:- disabling functionality;
- restricting access;
- stopping automated actions;
- isolating systems;
- disabling integrations;
- reverting to a previous model;
- increasing human oversight; or
- suspending the AI system.
17. Safety Protection
Where an AI incident may cause physical or significant individual harm, protection of affected people should take priority. Actions may include:- stopping affected operations;
- increasing human intervention;
- restricting system capabilities;
- initiating emergency procedures; and
- notifying appropriate authorities or responsible personnel.
18. Security Containment
Security-related AI incidents should be coordinated with applicable security incident processes. Actions may include:- credential revocation;
- access restriction;
- isolation;
- forensic preservation;
- malicious activity blocking; and
- security monitoring escalation.
19. Privacy Containment
Privacy-related incidents should be coordinated with applicable privacy incident processes. Actions may include:- restricting data access;
- stopping processing;
- securing exposed information;
- identifying affected data;
- preserving evidence; and
- initiating required notification assessment.
20. Incident Investigation
Material incidents should be investigated to determine:- what occurred;
- when it occurred;
- how it occurred;
- systems involved;
- data involved;
- people affected;
- controls involved;
- contributing factors; and
- potential root causes.
21. Evidence Preservation
Evidence should be preserved according to applicable requirements. Evidence may include:- logs;
- model versions;
- prompts;
- outputs;
- configurations;
- access records;
- system events;
- monitoring data;
- communications; and
- relevant documentation.
22. AI-Specific Evidence
Where relevant, incident investigation should preserve AI-specific information such as:- model identifier;
- model version;
- system prompt;
- user prompt;
- generated output;
- tool calls;
- agent actions;
- model configuration;
- retrieved data;
- relevant context;
- safety controls; and
- human interventions.
23. Root Cause Analysis
Material incidents should undergo root cause analysis where appropriate. Potential root causes may include:- model behavior;
- data quality;
- configuration;
- software defect;
- control failure;
- human error;
- inadequate monitoring;
- third-party failure;
- process weakness; or
- governance failure.
24. Contributing Factors
The investigation should distinguish between:- root cause;
- contributing factors;
- triggering event; and
- underlying governance weaknesses.
25. Control Failure Analysis
Where controls failed or were bypassed, the investigation should determine:- which control failed;
- whether the control was designed appropriately;
- whether it was implemented;
- whether it operated correctly;
- whether evidence existed; and
- whether compensating controls were effective.
26. Risk Reassessment
An incident should trigger risk reassessment where it indicates a material change in the AI system’s risk profile. Risk reassessment should consider:- likelihood;
- impact;
- affected populations;
- existing controls;
- residual risk; and
- newly identified risks.
27. Classification Reassessment
The AI system’s governance classification should be reviewed where the incident indicates that its existing classification may no longer be appropriate.28. Containment Decision
The incident owner should determine whether the AI system should:- continue operating;
- operate with restrictions;
- operate under enhanced monitoring;
- revert to a previous version;
- be temporarily suspended; or
- be permanently withdrawn.
29. Business Continuity
Where an AI system is operationally critical, incident response should consider continuity requirements. Alternative measures may include:- manual processes;
- fallback systems;
- alternative providers;
- reduced functionality; and
- temporary suspension of affected services.
30. Communication
Incident communication should be proportionate to severity. Potential stakeholders include:- affected users;
- business owners;
- executives;
- security;
- privacy;
- legal;
- regulators;
- customers;
- suppliers; and
- other affected parties.
31. Regulatory Notification
Where required, the organization should assess whether an incident triggers notification or reporting obligations. The assessment should be performed by appropriately authorized personnel.32. External Communication
External communications should be coordinated through authorized functions. Communications should be:- accurate;
- timely;
- proportionate;
- consistent; and
- appropriately documented.
33. Incident Remediation
Corrective actions should address identified weaknesses. Remediation may include:- model changes;
- control improvements;
- configuration changes;
- additional testing;
- monitoring improvements;
- training;
- process changes; and
- governance changes.
34. Corrective Action Plan
Material corrective actions should identify:- finding;
- required action;
- owner;
- priority;
- target date;
- evidence;
- verification; and
- closure criteria.
35. Preventive Actions
Where appropriate, preventive actions should be implemented to reduce recurrence. Preventive actions may address:- system design;
- controls;
- testing;
- monitoring;
- training;
- procedures;
- governance; and
- supplier management.
36. Incident Closure
An incident should only be closed when:- containment is complete;
- required investigation is complete;
- immediate risks are addressed;
- required notifications are completed;
- corrective actions are assigned;
- residual risk is understood; and
- closure is authorized.
37. Incident Review
Material incidents should receive a formal post-incident review. The review should consider:- response effectiveness;
- containment;
- communication;
- decision-making;
- controls;
- monitoring;
- root cause;
- remediation; and
- lessons learned.
38. Lessons Learned
Lessons learned should be incorporated into the AIGO governance system. Potential updates include:- controls;
- risk assessments;
- system profiles;
- procedures;
- monitoring;
- training;
- architecture;
- testing; and
- governance requirements.
39. Incident Records
Incident records should be maintained according to applicable retention requirements. Records should include:- incident identifier;
- date and time;
- AI system;
- severity;
- description;
- impact;
- containment;
- investigation;
- evidence;
- root cause;
- actions;
- notifications;
- approvals; and
- closure.
40. Incident Traceability
Incident records should maintain traceability to relevant governance artifacts. The organization should be able to demonstrate: Incident → AI System → Risk → Controls → Response → Investigation → Remediation → Lessons Learned41. Incident Monitoring
Open incidents should be actively monitored until closure. Monitoring should track:- actions;
- owners;
- deadlines;
- severity;
- residual risk;
- outstanding decisions; and
- escalation requirements.
42. Incident Escalation
Incidents should be escalated when:- severity increases;
- containment fails;
- new affected parties are identified;
- regulatory exposure increases;
- risk exceeds tolerance;
- required actions are delayed; or
- the incident becomes a major or critical event.
43. Incident Metrics
Organizations may establish incident management metrics. Examples include:- incident count;
- incident severity;
- time to detection;
- time to containment;
- time to resolution;
- repeat incidents;
- control-related incidents;
- unauthorized AI actions; and
- overdue corrective actions.
44. Incident Trends
Incident trends should be periodically reviewed to identify recurring weaknesses. Trend analysis may identify:- recurring model failures;
- repeated control failures;
- repeated security issues;
- recurring data problems;
- supplier issues; and
- systemic governance weaknesses.
45. Third-Party Incidents
Third-party AI incidents should be managed according to contractual and organizational requirements. The organization should determine:- provider notification;
- internal escalation;
- customer impact;
- regulatory implications;
- containment;
- provider remediation; and
- ongoing monitoring.
46. Incident and Change Management
Material incident remediation may require formal change management. Changes resulting from an incident should be assessed according to the AIGO change management procedure.47. Incident and Risk Management
Incident findings should feed into AI risk management. Where an incident demonstrates that an existing risk assessment is inaccurate or incomplete, the relevant risk record should be updated.48. Incident and Control Management
Incident findings should feed into control assessment. Controls that failed, were insufficient, or were bypassed should be reviewed and improved.49. Incident and Monitoring
Incident findings should be used to improve monitoring. Where appropriate, new:- indicators;
- thresholds;
- alerts;
- detection mechanisms; and
- review requirements
50. Responsibilities
Incident Owner- coordinate incident response;
- maintain the incident record;
- coordinate investigation;
- track actions;
- escalate appropriately; and
- confirm closure.
- provide system information;
- support containment;
- coordinate technical remediation;
- reassess system risk; and
- implement required corrective actions.
- oversee governance implications;
- coordinate escalation;
- maintain governance records;
- monitor material incidents; and
- report significant incidents.
- provide security, privacy, legal, compliance, risk, technical, safety, communications, or other specialist support as required.
51. Incident Management Workflow
The standard workflow should be:- Detect or receive incident report.
- Register incident.
- Perform initial triage.
- Determine severity.
- Assign incident owner.
- Initiate containment.
- Escalate where required.
- Preserve evidence.
- Investigate incident.
- Determine root cause and contributing factors.
- Assess risk and control implications.
- Determine notification requirements.
- Implement remediation.
- Validate remediation.
- Conduct post-incident review.
- Update governance artifacts.
- Record lessons learned.
- Obtain closure approval.
- Close incident.
52. Continuous Improvement
The incident management process should be improved based on:- incidents;
- near misses;
- response experience;
- assurance findings;
- audit findings;
- monitoring results;
- stakeholder feedback;
- regulatory developments; and
- changes in AI technology.
53. Procedure Review
This procedure should be reviewed periodically and when material changes occur. Review triggers may include:- significant incidents;
- changes to AIGO requirements;
- regulatory developments;
- assurance findings;
- changes to organizational incident management;
- changes in AI technology; and
- implementation experience.
54. Procedure Status
Document: AIGO AI Incident Management Procedure Version: 0.1 Status: Draft Working Name: AIGO Full Name: AI Governance Operating Framework Document Identifier:AIGO-PROC-008
Document Type: Operational Procedure
This procedure establishes the operational process for identifying, managing, investigating, resolving, and learning from AI incidents throughout the AI governance lifecycle.
