…

LEAD ENGINEER - Cloud Management

Happiest Minds Technologies · IT Services & Consulting

  • Bengaluru
  • On-site
  • Posted today
  • Software Engineering

About the job

Cloud Platform Observability Engineer

Location: Bengaluru, India

Years of Experience: 7+ years

Job Summary: We are seeking two Cloud Platform Observability Engineers to enhance our understanding of the health and behavior of mission-critical services operated by the Cloud Foundation team. This role focuses on the services themselves, discovering key signals, assisting service owners in exposing them safely, and transforming telemetry into actionable insights for operators and stakeholders. The services in scope include HashiCorp Vault, an internally developed key orchestration service built with Python and Django, Argo CD, bespoke Prometheus exporters, and ongoing OpenTelemetry experiments. The successful candidates will connect service architecture and operational needs, advancing a consistent approach to metrics, logs, traces, service-level indicators, dashboards, and alerts.

Responsibilities:

Build a service-centric observability strategy for mission-critical platforms, prioritizing telemetry based on user impact, operational risk, and diagnostic value.
Collaborate with service and infrastructure owners to understand architecture, dependencies, critical workflows, failure modes, and the telemetry needs of operators.
Define practical instrumentation standards and telemetry conventions for metrics, structured logs, traces, resource attributes, naming, and ownership.
Enhance application-level instrumentation for the Python/Django key orchestration service, ensuring useful request, dependency, and workflow signals while protecting sensitive information.
Configure and improve Vault and Argo CD telemetry, working with service owners to expose meaningful health, workload, and dependency signals.
Develop, test, document, and maintain bespoke Prometheus exporters, ensuring metrics are accurate, stable, and safe to operate.
Advance OpenTelemetry adoption through measured experiments and reusable patterns for instrumentation, collection, processing, and export.
Identify service-level indicators and assist in defining service-level objectives that reflect customer-facing reliability.
Design and refine dashboards and alerts to help operators recognize impact, isolate causes, and take clear next steps, while reducing noisy notifications.
Correlate telemetry with deployments, configuration changes, dependencies, and incidents to reveal regressions and emerging risks.
Establish guardrails for telemetry quality, including cardinality, sampling, retention, and data access.
Partner with Cloud Foundation and development teams to translate observability findings into instrumentation improvements and actionable runbooks.
Share findings and recommendations with technical and non-technical stakeholders, improving documentation for service behavior and operational diagnosis.

Mandatory Skills:

Experience designing or improving observability for production services, with a strong understanding of telemetry's role in reliability and troubleshooting.
Hands-on experience with Prometheus metrics and exporters, including metric design and diagnosing data-quality issues.
Practical experience with OpenTelemetry concepts and at least one of its signals: traces, metrics, or logs.
Proficiency in Python and experience with web application instrumentation; Django experience is a plus.
Strong understanding of distributed systems, service dependencies, APIs, and common failure patterns.
Experience building dashboards and actionable alerts with an observability or monitoring platform.
Familiarity with Kubernetes and GitOps workflows; experience with Argo CD is advantageous.
Ability to work across application and infrastructure layers, communicate trade-offs, and make incremental improvements in production environments.
Understanding of security and privacy considerations for telemetry, including preventing sensitive data capture.
Strong written communication skills for documenting instrumentation, dashboards, alerts, and operational findings.

Preferred Skills:

Experience operating or instrumenting HashiCorp Vault or other security-sensitive infrastructure services.
Experience with service-level indicators, service-level objectives, and reliability measurement practices.
Experience with OpenTelemetry Collector pipelines and telemetry backends.
Experience with GCP or Azure, Terraform, and infrastructure-as-code practices.
Experience analyzing production performance and reliability data to identify patterns and recommend changes.

Deal Breaker Skill: Cloud Management