This curriculum spans the equivalent depth and structure of a multi-workshop operational resilience program, integrating technical, procedural, and cross-functional responses to power outages across IT and facilities domains.
Module 1: Establishing Outage Detection and Alerting Infrastructure
- Configure redundant monitoring agents across geographically distributed data centers to ensure detection during regional power failures.
- Integrate UPS and PDU telemetry into monitoring systems using SNMP or vendor-specific APIs for real-time power status visibility.
- Define alert thresholds for battery runtime remaining on critical systems to trigger escalation before shutdown.
- Implement alert suppression rules during planned maintenance to prevent alert fatigue while preserving anomaly detection.
- Design alert routing logic that prioritizes on-call personnel based on location and system ownership during multi-site outages.
- Validate failover of monitoring systems to backup power and secondary sites through scheduled blackout simulations.
Module 2: Power Infrastructure Dependency Mapping
- Document physical server-to-PDU, PDU-to-UPS, and UPS-to-generator dependencies using asset management databases.
- Identify single points of power failure by auditing rack-level PDU diversity across separate UPS circuits.
- Map virtualized workloads to underlying physical hosts and correlate with power zones for impact analysis.
- Integrate building management system (BMS) data with IT service maps to visualize cross-domain dependencies.
- Classify systems by power criticality (e.g., Tier 0 for core network, Tier 1 for application servers) to prioritize response.
- Update dependency diagrams automatically using CMDB integrations with power distribution monitoring tools.
Module 3: Root-Cause Analysis Methodology for Power Events
- Apply the 5 Whys technique to distinguish between immediate cause (e.g., circuit tripped) and systemic cause (e.g., lack of load balancing).
- Correlate timestamps from UPS logs, generator start records, and server shutdown events to reconstruct outage sequence.
- Use fault tree analysis to model cascading failures from power loss to application unavailability.
- Preserve volatile data from UPS control panels and PDU event logs before system reset or reboot.
- Interview facilities personnel to determine whether manual interventions (e.g., bypass switch activation) contributed to failure propagation.
- Compare actual load draw against designed circuit capacity to assess overprovisioning or design flaws.
Module 4: Operational Response During Power Outages
Module 5: Post-Outage Forensics and Evidence Preservation
- Collect UPS event logs, including input voltage, load percentage, and battery health metrics prior to failure.
- Extract Windows Event Viewer or Linux kernel logs showing ACPI power state transitions and shutdown initiation.
- Preserve generator runtime logs and fuel level records to assess performance during the event.
- Archive network device logs indicating loss of power to PoE devices such as phones and access points.
- Obtain utility provider incident reports for grid-level disturbances affecting the facility.
- Secure physical access logs to electrical rooms to rule out unauthorized human interaction.
Module 6: Cross-Team Coordination and Escalation Protocols
- Define clear escalation paths between IT operations, facilities management, and external power vendors.
- Conduct joint tabletop exercises with facilities teams to align on outage response roles and communication channels.
- Establish shared incident bridges with standardized naming conventions for power-related events.
- Integrate facilities staff into incident management platforms for real-time status updates and task assignment.
- Document mutual dependencies in runbooks, such as IT requiring facilities to reset breakers before system restart.
- Implement joint post-mortem reviews to reconcile IT system failures with physical power event timelines.
Module 7: Mitigation Strategy Development and Validation
- Redesign PDU circuits to eliminate shared breakers across redundant server pairs in high-availability clusters.
- Upgrade UPS battery packs based on age, runtime testing results, and projected load increases.
- Install additional transfer switches to reduce single points of failure in generator-to-rack paths.
- Implement dynamic load shedding rules that automatically power down non-essential systems during brownouts.
- Conduct annual failover tests that simulate full utility loss to validate generator and UPS performance under load.
- Revise capacity planning models to account for increased power density from newer hardware deployments.
Module 8: Regulatory Compliance and Documentation Standards
- Align power continuity practices with ISO 22301 requirements for business continuity management.
- Maintain audit trails of UPS maintenance records, battery replacement logs, and generator test results.
- Classify power-related incidents according to internal severity levels for consistent reporting and trending.
- Document power architecture changes in network diagrams with version control and change request references.
- Ensure outage reports include root cause, impact duration, and remediation steps for regulatory review.
- Validate that SLAs with colocation providers specify uptime guarantees and compensation terms for power failures.