…

CONSULTANT

Atos · IT Services & Consulting

  • Chennai, India
  • On-site
  • Posted today
  • Apply by 30 Oct
  • IT & Infrastructure

About the job

skillset : Java , Observability (ELF , Grafan , Splunk) , Github , Service now(ticketing) , AWS/GCP

Replacement of Harikaran.

Roles & Responsibilities:

Provide hands-on support for the runtime operation of our applications, ensuring high availability and performance.
Collaborate with software engineering and infrastructure teams to troubleshoot and resolve runtime issues, including performance bottlenecks, scalability challenges, and system failures.
Contribute to the design and implementation of monitoring, alerting, and logging solutions to proactively identify and address potential runtime issues.
Participate in incident response and root cause analysis efforts to ensure the stability and resilience of the applications.
Work closely with cross-functional teams to understand application requirements and provide input on runtime and operational considerations during the software development lifecycle.
Contribute to the development and maintenance of runtime automation and tooling to streamline operational processes and improve efficiency.
Develop common framework components (to be leveraged by enterprise applications), define standards for configuration, monitoring, reliability, and performance engineering
Create automation and ensure automated tests are completed for new features.
Good attitude, communication, willingness to learn and collaborate.
Continuously improve automated remediation tasks to ensure the highest levels of availability.
Cloud: Manage secure, scalable, and highly available cloud infrastructure.
Kubernetes & Containers: Deploy, operate, and troubleshoot containerized workloads.
Observability: Implement monitoring, logging, tracing, dashboards, and actionable alerts.
Reliability: Define SLOs/SLIs, manage error budgets, and improve service availability.
Networking: Troubleshoot DNS, TCP/IP, HTTP/S, TLS, routing, and load balancing.
Linux & Systems: Administer and troubleshoot Linux systems and performance issues.
Programming/Scripting: Automate operational tasks using Python, Go, or Shell.
Infrastructure as Code: Provision and manage infrastructure using Terraform or equivalent IaC tools.
CI/CD: Build and maintain automated, reliable deployment pipelines.
Incident Management: Respond to incidents, perform RCA, and implement preventive actions.