•Lead major-incident response as Manager on Duty by coordinating triage, technical bridge calls, stakeholder communications, escalations, restoration activities, and post-incident follow-up in a 24/7 environment.
•Govern change intake, risk assessment, scheduling, readiness, and Change Advisory Board execution to improve change quality, reduce change-related incidents, and support stable releases.
•Strengthen Command Center operations and service resilience by improving runbooks, monitoring and escalation practices, leading problem reviews and root-cause follow-up, supporting Disaster Recovery exercises, and producing operational dashboards and executive updates.
•Partner with platform teams to enhance ServiceNow ITSM processes, maintain CMDB data quality, and improve integrations across monitoring, event-management, and automation tools.
•Coordinate managed service providers and technology vendors, mentor peers, and influence cross-functional teams to implement corrective actions and continuous-improvement initiatives.
•Major Incident Management and MOD duties: Act as Manager on Duty for high-impact events; coordinate triage, communications, and technical bridge calls; make time-sensitive decisions; drive service restoration and post-incident follow-ups *
•Change Management and CAB execution: Govern change intake, risk assessment, scheduling, and readiness; lead/participate in CAB; enforce standards; track change outcomes and reduce change-related incidents *
•Command Center operations and monitoring: Oversee 24/7 Command Center runbooks, alert handling, on-call rotations; ensure effective monitoring thresholds, event correlation, and escalation paths *
•Continuous improvement, Problem Management, and root cause analysis: Lead problem reviews; ensure root cause identification; implement corrective/preventive actions; mature operational best practices *
•Tooling and platform administration: Partner with platform engineering to configure/enhance ServiceNow ITSM modules; maintain CMDB data quality; integrate monitoring/event management tools and automation
•Reporting, analytics, and operational health: Produce dashboards for incident, change, and problem metrics; provide executive summaries; identify trends and risks; recommend improvement initiatives
•Vendor and stakeholder management: Coordinate with MSPs and SaaS vendors for escalations and SLAs; align expectations with business stakeholders; ensure clear communications during events
•Disaster Recovery support: Support DR planning, exercises, and after-action reviews; validate operational readiness; ensure Command Center integration into DR events
About You
•5–7 years of experience in IT operations or support, with demonstrated leadership across Incident, Problem, and Change Management in production environments.
•Experience leading major incidents as a primary liaison and participating in Manager on Duty or equivalent on-call rotations for 24/7 operations.
•Strong knowledge of ITIL and ITSM practices, change governance, Change Advisory Board execution, technical troubleshooting, root-cause analysis, and risk-based decision-making.
•Hands-on experience with ServiceNow ITSM, CMDB management, operational dashboards, and collaboration across network, server, cloud, application, security, and vendor teams.
•Clear executive, technical, and business communication skills, with the ability to coordinate teams and make time-sensitive decisions during high-impact events.
•Bachelor’s degree in Information Technology, Computer Science, Engineering, a related technical field, or equivalent practical experience.