•Drive Cloud Migration: Drive the end-to-end migration of production workloads from on-premise data centers to GCP.
•Architect for Reliability: Design and implement production-grade Kubernetes environments and GCP architectures that prioritize 99.99+% availability.
•Operational Excellence: Improve incident triage and improved recovery timeby implementing cloud-aware diagnostics and automated recovery patterns.
•Empower Developers: Reduce friction in the development lifecycle by optimizing platform response times and ensuring reliable, repeatable deployment events.
•Platform Standards, Security, and Governance: Embed security and compliance controls directly into platform abstractions (policy-as-code, identity, networking) and partner with security and compliance teams to meet regulatory requirements by design.
•Mentor and Standardize: Build patterns and shared tooling that allow the broader engineering teams to operate confidently in the cloud without external reliance.
12-Month Definition of Success
Within 12 months, you will have demonstrably delivered:
•Cloud Migration Execution
•Successful migration of production workloads to public cloud
•Prevented customer-impacting regressions attributable to platform design
•Clear, documented cloud architectures and operational runbooks
•Reliability & Operations
•Sustained 99.99%+ availability during and after migration
•Improved incident triage and improved recovery time through cloud-aware diagnostics
•Production-grade Kubernetes operations in cloud
•Team Uplift
•Platform team operating independently and confidently in cloud
•Shared standards, tooling, and golden paths adopted across teams
•Reduced reliance on external consultants or reactively addressing infrastructure issues
•Developer Productivity
•Faster platform response times for developer issues
•More reliable deployments and fewer rollback events
•Clear reduction in developer-reported friction related to infrastructure
Basic Qualifications
•10+ Total years of experience in a developer, devops, SRE, cloud role
•5+ years cloud experience (on-prem to cloud migration experience would be critical)
•5+ years Kubernetes and microservices operations experience at scale
•Exceptional incident triage and debugging skills in high-pressure environments
•Strong experience with infrastructure-as-code, CI/CD, and observability tooling.