…

Observability Engineer - Managed Services Engineer II

Deloitte USI (US-India offices) · Consulting & Professional Services

  • Hyderabad, Telangana, India
  • Hybrid
  • Posted 2 days ago
  • IT & Infrastructure

About the job

Artificial Intelligence & Engineering

AI & Engineering leverages cutting-edge engineering capabilities to help build, deploy, and operate integrated/verticalized sector solutions in software, data, AI, network, and hybrid cloud infrastructure. These insights are powered by engineering for business advantage, helping transform mission-critical operations.

Join our AI & Engineering team to help transform technology platforms, driving innovation, and help make a significant impact on our clients' achievements. You’ll work alongside talented professionals reimagining and re-engineering operations and processes that could be critical to businesses.

As a Managed Services Engineer II, you will design, build, deploy, and optimize enterprise-scale AI solutions that help clients automate workflows and improve business and information technology operations. In this role, you will work across the full delivery lifecycle, from requirements gathering and solution design through deployment, monitoring, and ongoing optimization. You will collaborate with cross-functional teams to translate business needs into scalable technical solutions and support high-quality delivery in production environments. You will also contribute to implementation planning, issue resolution, and continuous improvement initiatives.

Work you'll do

As a Managed Services Engineer II on the Hybrid Cloud Infrastructure team, you will be responsible for designing, building, deploying, and optimizing AI-powered agents and enterprise AI solutions that automate complex workflows and transform IT and business processes.

The Observability Platform Engineer is responsible for implementing, configuring, automating, and supporting enterprise observability solutions across hybrid cloud environments. This role works with architects, application teams, cloud engineers, and security stakeholders to deliver reliable monitoring, logging, tracing, and alerting using AppDynamics, ELK, OpenSearch, and DevOps automation tools.
Implement and maintain observability solutions using AppDynamics, Elasticsearch, Logstash, Kibana, and OpenSearch.
Configure dashboards, alerts, monitors, log pipelines, parsing rules, and search capabilities.
Onboard applications, infrastructure, databases, and cloud services into observability platforms.
Develop and maintain Infrastructure as Code using Terraform
Integrate monitoring and logging activities into GitHub pipelines.
Support observability deployments across AWS environments.
Configure and troubleshoot Docker, Kubernetes, and Helm-based workloads.
Monitor platform health, availability, performance, capacity, and data retention.
Investigate and resolve monitoring, logging, alerting, and integration issues.
Support migrations from legacy monitoring tools to ELK, OpenSearch, or AppDynamics.
Apply IAM, RBAC, SSL/TLS, access control, and compliance-monitoring practices.
Create operational documentation, runbooks, dashboards, and knowledge articles.
Participate in incident, problem, and change-management activities.
Collaborate with application, cloud, security, and DevOps teams to improve service reliability.
Contribute to observability standards, reusable automation, and platform best practices.

The team

At Hybrid Cloud Infrastructure we deliver solutions spanning Hybrid Cloud, Advanced Connectivity, AI Data Centers, High-Performance Computing, and AI Infrastructure to help clients achieve their desired outcomes. Our offerings include engineered transformation services for hybrid cloud infrastructure and platforms, prioritizing resiliency, optimization, and extensive automation. We integrate advanced connectivity with AI infrastructure and enterprise networks to boost operational efficiency and enable real-time data processing. Additionally, we provide comprehensive management of all facets of operations for hybrid cloud infrastructure and field operations.

Location: Bengaluru/Hyderabad

Shift Timings: 24/7 Rotational

Qualifications

Required:

3–5 years hands-on experience on the below
Strong expertise in AppDynamics (APM), ELK Stack (Elasticsearch, Logstash, Kibana), and OpenSearch.
Observability: New Relic, Datadog, Dynatrace, Prometheus, Grafana, OpenTelemetry, Jaeger, Distributed Tracing, Logs, Metrics, APM.
SRE: SLI, SLO, SLA, Error Budgets, Incident Management, Problem Management, Capacity Planning, Reliability Engineering, Automation, DR, Resilience Engineering
Deep understanding of DevOps tools – GitHub Actions, Docker, Kubernetes, Helm.
Experience in Infrastructure automation (Terraform).
Knowledge of cloud platforms – AWS.
Programming/scripting in Python, Shell, or Groovy.
Strong understanding of IAM, RBAC, SSL/TLS, and compliance monitoring.
Willingness to participate in on-call or production-support activities, where required.
ITSM: ServiceNow, ITIL, Incident, Problem, Change and Service Management
Excellent leadership, communication, and problem-solving skills.
BE/B.Tech/M.C.A./M.Sc (CS) degree or equivalent from accredited university.

Preferred:

AppDynamics Certified Implementation Professional
Elastic Certified Engineer / Observability Engineer or OpenSearch Practitioner
AWS Solutions Architect
Certified Kubernetes Administrator (CKA)
DevOps or SRE Certification

#HCIFY27