Developer III - DevOps Engineering
UST · IT Services & Consulting
- Trivandrum
- On-site
- Posted yesterday
- IT & Infrastructure
About the job
We are looking for a highly skilled DevOps Site Reliability Engineer (SRE) to design, implement, automate, and maintain highly available, scalable, and secure infrastructure platforms. The ideal candidate will have strong expertise in cloud technologies, CI/CD, infrastructure automation, monitoring, and incident management while ensuring system reliability and performance. Key Responsibilities Design, build, and maintain cloud-native infrastructure on AWS, Azure, or GCP. Develop and manage CI/CD pipelines for application deployment and release automation. Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or ARM templates. Manage and optimize Kubernetes clusters and containerized applications. Monitor system health, application performance, and infrastructure utilizing tools such as Prometheus, Grafana, Datadog, Splunk, and ELK Stack. Ensure platform reliability, scalability, availability, and security. Conduct root cause analysis (RCA) for production incidents and implement preventive measures. Automate operational tasks through scripting and orchestration tools. Collaborate with development, QA, security, and infrastructure teams to improve deployment processes. Define and maintain SLOs, SLIs, and SLAs. Support disaster recovery, backup, and business continuity initiatives. Implement security best practices and compliance requirements across environments.