Design and build the tooling layer
•Own the diagnostic tooling strategy — including AI-assisted tooling — so that establishing whether an issue is infrastructure or market data, and which clients it touches, takes seconds rather than a manual trawl through logs and jobs. Tooling, not documentation alone, is how we intend to get response times down.
•Write and review production-quality Python that eliminates toil: runbook automation, health checks, self-service diagnostics and remediation used across the team.
•Design automated recovery and self-healing for analytics jobs, so a single failed task does not become a missed client deliverable.
•Set the observability standard — design monitoring and alerting around business outcomes (late valuations, missing risk measures, stale prices) rather than host and container metrics, and drive alert noise down so pages are actionable and rare.
Own production and lead incidents
•Own incident response for client-impacting issues: command the incident, mitigate, communicate, and run blameless postmortems that end in fixes rather than findings.
•Investigate at the application level — logs, stack traces, job histories and the code itself — to establish root cause, then fix it or hand the development team a diagnosis precise enough to act on.
•Track service commitments and recurring pain: report on timeliness and failure trends, and push through the fixes that retire repeat issues.
•Set the boundary for first-line infrastructure work — stale hosts, memory pressure, connectivity failures, host loss — so routine cases are remediated locally and deeper platform work escalates cleanly.
Harden the batch and end-of-day architecture
•Own and harden large-scale batch and EOD processing: job orchestration, dependency management, retries, idempotency, backfills and safe reruns.
•Keep the overnight window on schedule — find where time is going across stages and drive the fixes with the development teams.
Own the data and market data interface
•Support the reference data, market data and timeseries stores behind pricing and risk — schema and index design, replication, retention and query performance in NoSQL and timeseries systems.
•Operate data quality and reconciliation controls with the market data team: completeness checks, stale-data and anomaly alerting, vendor feed failure handling.
•Debug data-shaped production problems — a missing fixing, a bad identifier mapping, a holiday calendar difference, a failed download job, a partial vendor file — traced through to the affected downstream analytics.
•Act as a single escalation path with the market data team on front-office-facing issues, sharing triage and root cause rather than passing tickets between functions.
Lead and influence
•Partner with quant developers and risk engineers to make reliability a design input rather than a post-launch concern.
•Mentor engineers on incident triage, investigation technique and on-call practice, and set the standards for automation and operational readiness.
•Contribute to production readiness reviews for new services and client onboardings.
•Support client escalations on risk and analytics output, and be the person client-facing teams want in the room during an incident.
•Document architecture, runbooks and failure modes.
What you'll bring
Required
•7+ years in software engineering for a business-critical platform, with real ownership of services.
•Strong Python , plus shell. You must be comfortable reading and debugging application code to find root cause — this role is not runbook execution.
•Demonstrated troubleshooting depth — working from a vague client symptom through logs, job histories, data and code to a precise diagnosis.
•Hands-on experience operating batch and scheduled workloads at scale , including workflow orchestration (Airflow, Prefect, Step Functions or equivalent).
•Solid grounding in distributed systems — you can reason about partial failure, retries, idempotency, consistency, and queue and backpressure behaviour under load.
•Working knowledge of AWS — enough to operate applications confidently (compute, S3, IAM, CloudWatch, basic networking) and to tell an infrastructure problem from an application one.
•Practical NoSQL experience (MongoDB, DynamoDB, Cassandra or similar): data modelling, indexing, replication and performance troubleshooting.
•Observability in practice — Prometheus, Grafana, Datadog, OpenTelemetry or equivalent — and sensible alert design.
•On-call experience with structured escalation and postmortem writing , and the judgement to lead an incident rather than only participate in one.
•A general understanding of financial systems and genuine appetite to learn the business. You should be able to hold a conversation about what a valuation or a risk number is for, and be willing to go deeper.
•Clear written and verbal communication, and the ability to work effectively in a globally distributed team.
•Degree (B.Tech / M.Tech / MCA / MSc) in computer science, engineering or a related field.