•Lead incident response, root cause analysis, and service restoration activities for business-critical planning and allocation platforms.
•Build and coach the India-based SRE team while supporting hiring, onboarding, capability development, and operational excellence.
•Establish observability, monitoring, logging, SLOs, SLIs, operational dashboards, and reliability measurements.
•Drive DevSecOps, automation, AI-assisted operations, self-healing workflows, and continuous improvement initiatives.
•Partner with SAP, SAS, Azure, middleware, engineering, and business teams to manage releases, risks, dependencies, and platform priorities.
About You
•8+ years of experience in Site Reliability Engineering, Production Support, Platform Engineering, DevOps, Cloud Operations, or Enterprise Application Operations.
•Strong understanding of reliability engineering, incident management, change management, observability, and operational readiness practices.
•Hands-on experience with Azure services, CI/CD, Infrastructure as Code, enterprise integrations, APIs, scripting, and monitoring platforms.
•Experience supporting retail, supply chain, planning, allocation, forecasting, ERP, or other business-critical enterprise platforms, with exposure to AppliedAI and AI-enabled operations.