…

SRE / Reliability Engineer (Lead)

Tenarai (formerly Infogain) · IT Services & Consulting

  • Bangalore, India
  • On-site
  • Posted today
  • IT & Infrastructure
  • Full-Time / Contract

About the job

&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt1. Key Responsibilities&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtNew Relic as code&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Provider modules — Build reusable Terraform modules with the New Relic provider for dashboards, alert policies, NRQL conditions, workflows, destinations (Jira, Teams), synthetics, cloud-integration links and entity tags.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Reverse-engineer click-ops — When New Relic SMEs or engineers build or modify dashboards/alerts in the UI (including tuned out-of-the-box dashboards), import and codify them so the state is recoverable.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Per-technology packs — Package each proven pattern (e.g., MySQL, Airflow, EC2, Vault) as a module other teams can consume with a few variables.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtApplication &ampamp; infrastructure codification&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Agent rollout automation — Update onboarding scripts, AMIs/container images and Terraform so the New Relic agent replaces Datadog on Meridian nodes and every new per-client instance comes up instrumented.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Cloud integration plumbing — Codify IAM roles, CloudWatch metric streams, Firehose and log subscriptions across the AWS account structure (AWS Control Tower / AFT), with Azure and GCP equivalents where needed.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Pipeline hooks — Add New Relic change tracking/deployment markers into Jenkins and GitHub Actions deploy pipelines; support ECS/Fargate (COAL) and Argo CD deployments.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtCI/CD, governance &ampamp; recovery&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• GitHub Actions workflows — Plan/apply pipelines with PR review, policy checks, drift detection and environment promotion for all observability code.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• State &ampamp; secrets — Remote state, workspace strategy per account/environment; New Relic API keys and credentials managed in HashiCorp Vault.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Disaster recovery drill — Demonstrate that a New Relic account configuration can be recreated from code.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Backstage alignment — Expose observability modules as golden-path templates alongside the new platform automation (Terraform, GitHub Actions, Backstage, Argo CD, AFT, Control Tower).&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtDocumentation &ampamp; enablement&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Module READMEs, usage examples and a contribution guide so Cyderes engineers can self-serve observability as part of the definition of done.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtEnvironment the engineer will work in:&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtService group Technologies&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtMeridian — Provision Docker, AWS ECR, AWS SSM, Jenkins CI/CD pipelines; SaaS Customer Manager, Update Manager (per-client provisioning)&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtMeridian — Ingress Tunnel-free and tunnel-dependent connector VMs (~800 connectors), tunnel proxy server, MySQL data cache on AWS EC2 (per-client containers)&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtMeridian — Process / Analyze Apache Airflow (scheduler/workers) with PostgreSQL backend, Action Manager; Python ML Merger; LLM VM (Ollama, IBM Granite)&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtService group Technologies&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtMeridian — Visualize / Store Multi-tenant Web UI and API; Cloudflare (WAF/CDN/DNS/DDoS, R2); WorkOS (SSO/SAML); MongoDB Atlas; AWS S3&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtMeridian MCP Airflow + Merger (produce) ? S3 Uploader (export) ? AWS Step Functions / Lambda staging (transform) ? Amazon Neptune snapshot, bulk load, validation (load) ? Bedrock AgentCore Gateway, Lambda, Claude, MCP clients (consume)&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtVault HashiCorp Vault (secrets storage and rotation; SaaS dependency for every service — agentless monitoring)&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtCOAL (Deploy) AWS ECS, AWS Fargate; frontend, backend and worker services&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtPlatform AWS Control Tower, AFT (Account Factory for Terraform) pipeline, Terraform, GitHub Actions, Argo CD, Backstage&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtIntegrations Jira (incidents, on-call schedules, paging), Microsoft Teams, public status page, customer portal, Cyderes MCP&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtClouds / legacy AWS (primary, growing under private pricing), Azure and GCP; legacy Prometheus, Thanos, Grafana, Loki, Datadog, Azure Monitor&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt3. Required Technical Skills&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtArea Expectation&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtTerraform (must) Modules, remote state, workspaces, import blocks, for_each/dynamic blocks, testing (terraform test / Terratest), version pinning&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtNew Relic provider (must) newrelic_one_dashboard, alert policies and NRQL conditions, workflows and destinations, synthetics, cloud-account linking, tagging&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtGitHub Actions (must) Reusable workflows, OIDC to AWS, environments and approvals, plan-on-PR/apply-on-merge&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtAWS (must) IAM, Organizations, Control Tower and AFT, CloudWatch metric streams, Kinesis Firehose, EC2, ECS/Fargate, Lambda, S3&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtPlatform tooling Argo CD, Backstage, Jenkins, Docker; HashiCorp Vault for secrets&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtMulti-cloud Terraform for Azure (Azure Monitor) and GCP (Cloud Logging/Monitoring) for cross-cloud workflows and migrations to AWS&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtScripting Python, Bash, HCL; NRQL literacy to review conditions and dashboards&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gtPolicy &ampamp; quality tflint, checkov/tfsec or OPA, pre-commit; semantic versioning of modules&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt4. Domain &ampamp; Functional Knowledge&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Observability-as-code — treating dashboards, alerts and notification routing as versioned, reviewable artifacts.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• GitOps and platform engineering — golden paths, self-service templates, landing-zone account structures.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Change and release governance in a security company — least-privilege IAM, auditable changes, secrets never in code.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Disaster-recovery thinking — reproducible configuration, drift detection, tool-failure scenarios.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Multi-cloud migration awareness — applications moving from GCP to AWS may need their telemetry wiring rebuilt.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt5. Experience &ampamp; Qualifications&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• 6–9 years in DevOps, platform or cloud engineering, with 4+ years writing production Terraform.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Hands-on experience with the New Relic Terraform provider (or Datadog/Grafana providers with clear transferability).&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Built and operated GitHub Actions-based infrastructure pipelines in a multi-account AWS organization.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Experience importing and codifying existing, manually built configuration.&lt/span&gt&lt/p&gt&ltp&gt&ltspan style=&quotcolor:rgb(0, 0, 0);&quot&gt• Bachelor&aposs degree in Computer Science/Engineering or equivalent experience.&lt/span&gt&lt/p&gt