Lead I - DevOps Engineering
UST · IT Services & Consulting
- Hyderabad, Bangalore
- Hybrid
- Posted today
- IT & Infrastructure
About the job
Build and maintain automation pipelines supporting Continuous Integration and Continuous Delivery using Harness or GitHub Actions. • Strong Expertise on building IAC using Terraform . • Provide critical support on AWS cloud environments • Integrate Datadog and create custom dashboards, s and log pipelines for continuous monitoring and observability. • AWS resource provisioning, management, architecture with a focus on cost and optimization • Lead projects & initiatives to completion to improve and streamline operational processes and maximize resources • Interface with other teams to resolve complex issues that have implications beyond your own area • Assess infrastructure and application vulnerabilities and take remediation actions as appropriate. • Change management and Incident management experience using JIRA and ServiceNow. • Collaborate with developers to streamline code deployments and environment configurations. • Work with multiple teams to implement security best practices across Containers, EC2 and Serverless workloads. Embed security guardrails into CI/CD pipelines including SAST, SCA, DAST, container image scanning and SBOM generation, with policy-based gating and fail-fast controls. • Build and secure container platforms on EKS or ECS, including image hardening, admission controls, runtime security and patching cadence. • Participate in on-call rotation, incident response, blameless post-mortems and periodic disaster recovery and resilience testing. • Implement anomaly detection, event correlation and noise reduction across metrics, logs and traces to reduce fatigue and improve MTTD/MTTR. • Build auto-remediation and self-healing workflows triggered by observability signals using EventBridge, Lambda, Step Functions and SSM Automation. • Leverage ML and GenAI capabilities, such as Amazon Bedrock and agentic or MCP-based tooling, for intelligent triage, root cause analysis assistance and operational knowledge retrieval. • Design and implement AIOps solutions to monitor, analyze, and optimize AWS cloud infrastructure and costs. Required Qualifications • 5+ years of hands-on experience with core AWS services including EC2, ASG, Security Groups, ALB/NLB/WAF, NACLs, Routing, Route 53, VPC networking, EC2 Image Builder, EKS, ECS, ECR, Lambda and SSM. • Experience in setting up, securing and operating Kubernetes clusters across cloud and hybrid platforms, including upgrades, scaling and troubleshooting. • Strong knowledge of containerisation and orchestration on EKS and ECS, covering image hardening, registry management and runtime security. • Strong expertise in Infrastructure as Code using reusable Terraform modules, remote state management, drift detection and automated plan/apply workflows. • Experience working with AWS Bedrock foundation models and SageMaker Studio, including prompt design, model selection and cost-aware inference. • Experience integrating AIOps tooling with AWS services such as CloudWatch, EventBridge, Lambda, S3, EC2, ECS and RDS to enable automated detection and response. • Hands on Experience on testing and orchestrating Disaster Recovery Strategies on AWS cloud. • Hands-on experience building and operating CI/CD pipelines using Harness, GitHub Actions or Jenkins, including job promotion, approvals and rollback strategies. • Practical experience with application and infrastructure security tooling covering SAST, SCA, DAST, container image scanning and SBOM generation using tools such as SonarQube, JFrog Xray, checkamrx and Wiz. • Proficiency in scripting and automation using Python, Bash or powershell with strong Git and branching-strategy discipline. • Hands-on experience with observability platforms such as Datadog, CloudWatch, Dynatrace or OpenTelemetry, covering metrics, logs, traces, synthetic monitoring and dashboarding. • Experience building auto-remediation and self-healing workflows using EventBridge, Lambda, Step Functions and SSM Automation. • Exposure to agentic and GenAI-assisted operations, including Bedrock AgentCore, MCP-based tooling, RAG patterns and AI-assisted triage or root cause analysis. • FinOps and cost optimisation experience covering tagging governance, rightsizing, Graviton migration, Spot adoption, Savings Plans coverage and cost anomaly detection. • Experience with incident, problem and change management using JIRA and ServiceNow, including on-call participation, post-incident reviews and DR/resilience testing. • Strong communication, stakeholder management and mentoring skills, with the ability to lead technical initiatives independently in an Agile environment.