•Observability Engineering
•Design, implement, and maintain enterprise-grade observability solutions across applications, infrastructure, cloud platforms, and services.
•Build and maintain monitoring, alerting, dashboards, service health views, and operational telemetry.
•Develop standardized observability patterns for metrics, logs, traces, events, and application performance monitoring .
•Implement observability solutions using New Relic, Grafana, and other industry-standard tools .
•Develop reusable dashboards, alerts, instrumentation patterns, and observability components.
•Establish observability standards and best practices across engineering teams.
•Continuously improve signal quality by reducing alert noise, false positives, and non-actionable alerts.
•New Relic Engineering
•Hands-on engineering experience with New Relic APM, Infrastructure Monitoring, Browser Monitoring, Synthetic Monitoring, Logs, Distributed Tracing, NRQL, Alerts, Workloads and Dashboards .
•Design and implement New Relic monitoring and alerting strategies for enterprise applications.
•Develop complex NRQL queries , alert conditions, dashboards, and operational views.
•Configure and optimize New Relic agents and integrations.
•Implement application and infrastructure instrumentation.
•Develop reusable New Relic configurations and automation using APIs/IaC where appropriate.
•Participate in New Relic platform governance, licensing optimization, and standardization.
•Evaluate and implement emerging New Relic capabilities to improve engineering productivity and reliability.
•Grafana & Visualization
•Build and maintain operational dashboards using Grafana .
•Integrate Grafana with multiple telemetry and data sources.
•Design effective dashboards for application health, infrastructure, SRE, NOC, and executive operational visibility.
•Develop visualization standards and reusable dashboard templates.
•Understand the difference between visualization, monitoring, alerting, and observability , and apply each appropriately.
•OpenTelemetry & Modern Observability
•Experience with OpenTelemetry and modern telemetry architectures.
•Implement and manage telemetry collection for metrics, logs, and traces.
•Understand distributed tracing and service dependency mapping.
•Work with telemetry pipelines, collectors, agents, exporters, and integrations.
•Experience with technologies such as Prometheus, Loki, Elastic, Splunk, Datadog, Dynatrace, AppDynamics, or similar observability platforms is desirable.
•Evaluate new observability technologies and recommend solutions based on scalability, cost, reliability, and engineering value.
•Cloud Engineering
Strong cloud engineering fundamentals are expected, including experience with one or more major cloud platforms:
•AWS
•Microsoft Azure
•Google Cloud Platform
Experience should include:
•Compute, networking, storage, databases, containers, and cloud-native services.
•Cloud monitoring and logging.
•IAM and security fundamentals.
•Infrastructure automation.
•Cloud-native architecture and operational best practices.
•Troubleshooting distributed cloud environments.
•Automation & Infrastructure as Code
•Automate repetitive observability and operational activities.
•Develop scripts and tools using Python, Bash, Go, or similar languages .
•Use Terraform / OpenTofu or equivalent Infrastructure as Code technologies.
•Build reusable automation for dashboards, alerts, instrumentation, integrations, and configuration management.
•Integrate observability capabilities into CI/CD pipelines.
•DevOps & CI/CD
•Experience with modern CI/CD practices and tools.
•Hands-on experience with GitHub Actions, Jenkins, GitLab CI, Azure DevOps, or similar platforms .
•Integrate observability and quality gates into deployment pipelines.
•Implement deployment markers and release health monitoring.
•Enable automated validation of application and infrastructure health following deployments.
•SRE & Reliability Engineering
•Apply SRE principles to improve system reliability and operational maturity.
•Define and monitor SLIs, SLOs, and error budgets .
•Participate in incident investigation and root-cause analysis.
•Develop proactive monitoring and reliability solutions.
•Identify reliability gaps and engineer solutions to eliminate recurring incidents.
•Support capacity, performance, availability, and resilience engineering.
Required Skills & Experience