•Build a service-centric observability strategy for mission-critical platforms, prioritizing telemetry based on user impact, operational risk, and diagnostic value.
•Collaborate with service and infrastructure owners to understand architecture, dependencies, critical workflows, failure modes, and the telemetry needs of operators.
•Define practical instrumentation standards and telemetry conventions for metrics, structured logs, traces, resource attributes, naming, and ownership.
•Enhance application-level instrumentation for the Python/Django key orchestration service, ensuring useful request, dependency, and workflow signals while protecting sensitive information.
•Configure and improve Vault and Argo CD telemetry, working with service owners to expose meaningful health, workload, and dependency signals.
•Develop, test, document, and maintain bespoke Prometheus exporters, ensuring metrics are accurate, stable, and safe to operate.
•Advance OpenTelemetry adoption through measured experiments and reusable patterns for instrumentation, collection, processing, and export.
•Identify service-level indicators and assist in defining service-level objectives that reflect customer-facing reliability.
•Design and refine dashboards and alerts to help operators recognize impact, isolate causes, and take clear next steps, while reducing noisy notifications.
•Correlate telemetry with deployments, configuration changes, dependencies, and incidents to reveal regressions and emerging risks.
•Establish guardrails for telemetry quality, including cardinality, sampling, retention, and data access.
•Partner with Cloud Foundation and development teams to translate observability findings into instrumentation improvements and actionable runbooks.
•Share findings and recommendations with technical and non-technical stakeholders, improving documentation for service behavior and operational diagnosis.
Mandatory Skills:
•Experience designing or improving observability for production services, with a strong understanding of telemetry's role in reliability and troubleshooting.
•Hands-on experience with Prometheus metrics and exporters, including metric design and diagnosing data-quality issues.
•Practical experience with OpenTelemetry concepts and at least one of its signals: traces, metrics, or logs.
•Proficiency in Python and experience with web application instrumentation; Django experience is a plus.
•Strong understanding of distributed systems, service dependencies, APIs, and common failure patterns.
•Experience building dashboards and actionable alerts with an observability or monitoring platform.
•Familiarity with Kubernetes and GitOps workflows; experience with Argo CD is advantageous.
•Ability to work across application and infrastructure layers, communicate trade-offs, and make incremental improvements in production environments.
•Understanding of security and privacy considerations for telemetry, including preventing sensitive data capture.
•Strong written communication skills for documenting instrumentation, dashboards, alerts, and operational findings.