CI/CD & GitOps :
•Design, build, maintain, and enhance complex CI/CD pipelines using Jenkins, GitHub Actions, Azure DevOps, or similar tools.
•Implement reusable pipeline libraries, templates, quality gates, security scanning, automated testing, packaging, and deployment workflows.
•Implement and support GitOps-based deployment workflows using Argo CD, Flux, or equivalent technologies.
•Maintain declarative application and platform configuration, automated reconciliation, environment promotion, drift detection, rollback, RBAC, and secrets integration.
•Implement deployment strategies such as blue/green, canary, progressive delivery, and automated rollback where appropriate.
•Troubleshoot complex CI/CD, GitOps, and deployment issues and implement permanent corrective solutions.
AWS & Cloud Infrastructure:
•Implement and improve scalable, secure, resilient, and cost-efficient infrastructure primarily in AWS.
•Work with AWS services including EC2, ECS, EKS, IAM, VPC, ALB/NLB, S3, RDS, Route 53, CloudWatch, SSM, Lambda, and related platform services.
•Troubleshoot complex infrastructure, networking, IAM, DNS, security, performance, and application connectivity issues.
•Implement established cloud architecture, security, tagging, governance, reliability, and operational standards.
•Contribute implementation expertise, technical options, and proof-of-concept validation during solution and design reviews.
Infrastructure as Code:
•Design, develop, test, and maintain reusable infrastructure using Terraform/OpenTofu.
•Create and enhance modular IaC frameworks used across multiple products, environments, and AWS accounts.
•Implement Terraform state management, testing, validation, versioning, and lifecycle practices.
•Integrate infrastructure provisioning with CI/CD and GitOps workflows.
•Troubleshoot complex Terraform plan/apply issues, state issues, drift, and dependency problems.
•Participate in and provide technical feedback during peer reviews of Infrastructure as Code changes.
Automation & Platform Engineering :
•Develop automation and platform tooling using Python, Go, PowerShell, Bash, TypeScript, Ansible, or equivalent technologies.
•Build reusable automation, APIs, utilities, templates, and workflows that reduce manual engineering and operational activities.
•Implement Internal Developer Platform capabilities and self-service workflows based on established platform architecture and standards.
•Build and improve paved roads/golden paths that simplify infrastructure provisioning and application delivery.
•Maintain and improve artifact repository, configuration-management, and engineering automation capabilities.
Containers & Kubernetes:
•Build, configure, operate, and improve Kubernetes platforms, particularly Amazon EKS and Azure AKS where applicable.
•Implement standards for cluster configuration, networking, workload isolation, identity, security, scaling, storage, observability, and lifecycle management.
•Build and maintain reusable Helm charts and Kubernetes deployment patterns.
•Troubleshoot complex Kubernetes, container, networking, storage, identity, resource, and application deployment issues.
•Automate Kubernetes provisioning, upgrades, configuration, and operational activities.
Monitoring, Reliability & Troubleshooting :
•Implement monitoring, logging, metrics, tracing, dashboards, and alerting using Datadog, Grafana, Prometheus, CloudWatch, or similar platforms.
•Lead troubleshooting and root-cause analysis for complex infrastructure, platform, CI/CD, and deployment issues.
•Implement high-availability, resiliency, disaster-recovery, scalability, and performance patterns defined for platform workloads.
•Identify recurring operational issues and implement automation or permanent corrective actions.
•Create and improve operational runbooks, health checks, recovery procedures, and platform documentation.
•Participate in an on-call rotation where applicable.
AI & Engineering Automation:
•Apply AI and agentic capabilities to practical DevOps and Platform Engineering use cases, including automation, CI/CD, troubleshooting, incident analysis, documentation, and developer self-service.
•Build or integrate AI-enabled workflows using approved enterprise AI services such as Amazon Bedrock, Azure OpenAI/OpenAI, or equivalent platforms.
•Implement secure integrations between AI capabilities and approved cloud services, APIs, repositories, and engineering tools.
•Apply appropriate identity, authorization, observability, guardrails, cost controls, and human-in-the-loop mechanisms to AI-enabled automation.
Technical Leadership & Developer Enablement:
•Provide technical guidance within Platform Engineering initiatives and lead defined technical workstreams when required.
•Mentor engineers and share expertise across cloud, DevOps, GitOps, Kubernetes, automation, and platform engineering practices.
•Participate in design, code, infrastructure, and operational-readiness reviews and provide actionable technical feedback.
•Collaborate with architects and senior engineers to evaluate implementation approaches and architectural trade-offs.
•Partner with Engineering, Platform, CloudOps, SRE, Security, and other teams to deliver reliable platform capabilities.
Required Skills & Experience:
•7+ years of experience in DevOps, Cloud Engineering, Platform Engineering, SRE, Infrastructure Engineering, Software Engineering, or a related role.
•Strong hands-on experience with AWS and enterprise cloud infrastructure.
•Advanced experience with Terraform/OpenTofu and Infrastructure as Code, including reusable modules and multi-environment implementations.
•Strong experience designing, building, and troubleshooting CI/CD pipelines using Jenkins, GitHub Actions, Azure DevOps, or equivalent platforms.
•Strong hands-on experience with GitOps deployment tools such as Argo CD, Flux, or equivalent technologies.
•Strong hands-on experience with Docker, Kubernetes, EKS and/or AKS, and Helm.
•Experience with configuration management and automation using Ansible or equivalent technologies.
•Strong scripting/programming skills using one or more of Python, Go, PowerShell, Bash, TypeScript, or equivalent languages.
•Experience with Git and pull-request-based development workflows.
•Strong understanding of IAM/RBAC, cloud security, networking, DNS, load balancing, VPCs/subnets, routing, and security groups/firewalls.
•Experience with observability platforms such as Datadog, Grafana, Prometheus, CloudWatch, or equivalent tools.
•Ability to independently troubleshoot complex infrastructure, Kubernetes, CI/CD, GitOps, and deployment issues.
•Demonstrated experience mentoring engineers and providing technical guidance.