Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities:
•Design and Maintain Cloud Infrastructure - Architect, deploy, and support highly available, secure, and scalable cloud infrastructure for the ECP
•Drive CI/CD Automation - Develop, enhance, and maintain automated build, test, and deployment pipelines to improve release velocity, quality, and reliability
•Manage Containerized Platforms - Administer and optimize container-based workloads using ECS/Fargate and related technologies, ensuring performance, scalability, and resilience
•Lead Infrastructure as Code (IaC) Initiatives - Implement and maintain infrastructure provisioning through Terraform, CloudFormation, or equivalent IaC frameworks to ensure consistency across environments
•Ensure Platform Reliability and Operational Excellence - Monitor system health, performance, and availability; proactively identify risks and implement solutions to meet uptime and resiliency objectives
•Support Production Operations and Incident Management - Troubleshoot critical production issues, participate in on-call support, conduct root-cause analyses, and drive preventive actions to reduce recurring incidents
•Implement Security and Compliance Controls - Enforce security best practices, manage access controls, certificates, secrets, and support compliance initiatives aligned with enterprise and regulatory standards
•Collaborate Across Engineering Teams - Partner with application developers, architects, security teams, and platform stakeholders to deliver reliable infrastructure solutions, optimize operational processes, and mentor team members on DevOps best practices
•Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so
Required Qualifications:
•Bachelor's or master's degree in computer science, Engineering, Information Technology, or a related field
•7+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or Infrastructure Engineering
•5+ years of hands-on experience designing, deploying, and managing AWS cloud infrastructure
•Experience building and maintaining CI/CD pipelines using GitHub Actions, Jenkins, Azure DevOps, GitLab CI/CD, or equivalent platforms
•Experience with containerization technologies such as Docker and container orchestration platforms including ECS or Kubernetes
•Solid expertise in AWS services including ECS/Fargate, EC2, VPC, IAM, ALB, Route 53, S3, CloudWatch, Lambda, RDS, and CloudFormation/Terraform
•Proficiency in Infrastructure as Code (IaC) using Terraform, AWS CloudFormation, or similar technologies
•Proven solid scripting and automation skills using Python, Shell, Bash, or PowerShell
•AI & Automation Qualifications - Experience utilizing AI-assisted engineering tools such as GitHub Copilot, Amazon Q, ChatGPT, Claude, Cursor, or equivalent solutions
•Experience building AI-driven operational automation, monitoring enhancements, or self-healing infrastructure solutions is preferred
•Working knowledge of Generative AI technologies, Large Language Models (LLMs), and AI-powered developer productivity tools
•Understanding of AI governance, security, responsible AI practices, and enterprise adoption considerations
•Demonstrated ability to leverage AI tools to automate infrastructure management, troubleshooting, code reviews, documentation generation, and operational workflows
•Technical Competencies: - Experience implementing DevSecOps practices and integrating security scanning into CI/CD pipelines
•Experience supporting production-critical applications with defined SLA/SLO objectives
•Expertise in monitoring, observability, and logging platforms such as CloudWatch, Splunk, Datadog, ELK, Grafana, or Prometheus
•Solid understanding of cloud networking, security, identity management, and compliance controls
•Knowledge of high-availability architectures, disaster recovery, backup strategies, and zero-downtime deployment patterns
•Professional Skills - Experience mentoring engineers and promoting DevOps, automation, and operational excellence best practices
•Demonstrated ability to lead incident response, root cause analysis, and continuous improvement initiatives
•Demonstrated ability to collaborate effectively with developers, architects, security teams, and business stakeholders
•Proven solid analytical and problem-solving capabilities with a proactive operational mindset
•Proven excellent verbal and written communication skills