•Continuously discover cost and efficiency opportunities through production profiling, telemetry, cost data, capacity trends, and system-level analysis — across algorithmic inefficiencies, resource-heavy code paths, and architectural decisions — and turn ambiguous problems into prioritized engineering initiatives.
•Apply performance and capacity engineering techniques to understand CPU, memory, storage, network, and I/O behavior under real production workloads, and optimize the resulting resource footprint.
•Write production-grade code to implement the optimizations you identify — JVM/GC tuning, algorithmic and resource-efficiency improvements, re-architecting inefficient services — in systems that process petabytes of data daily.
•Define and track engineering efficiency metrics such as cost per GB ingested, cost per query, cost per event, resource utilization, or cost per customer workload, and translate the work into measurable impact for engineering and business stakeholders.
•Lead complex, cross-team engineering initiatives from problem discovery through design, implementation, rollout, and measurement — influencing teams where you don't have direct ownership — and help establish engineering patterns and practices that make cost and efficiency a continuous part of the development lifecycle.
•Partner with engineering teams in your product area to prioritize changes, and with developer infrastructure and Global SRE to align with the broader reliability roadmap.
•Participate in the SRE responsibilities for different product areas — SLOs, on-call, incident response, and blameless RCA — using those experiences to identify systemic reliability, performance, and efficiency improvements.
What you’ll have
•B.Tech, M.Tech, or equivalent degree in Computer Science or a related discipline.
•8+ years of industry experience with a demonstrated track record of ownership.
•Strong CS fundamentals — comfortable with algorithmic complexity, data-structure performance characteristics, and system design at scale.
•Ability to author production-ready code in at least one OO/systems language (Java, Scala, Go, C++, or similar) — depth of engineering ability matters more than which language.
•Experience with distributed systems and microservice architectures in production.
•Demonstrated track record of independently identifying ambiguous performance, scalability, or cost problems and driving engineering changes that produced measurable improvements in production.
•Strong ability to reason quantitatively about system behavior, capacity, performance, and cost, and use production data to validate hypotheses and measure outcomes.
•Working fluency with cloud infrastructure (AWS compute, storage, networking) — enough to reason about cost and architectural tradeoffs.
•Comfort moving across the stack, from application code to the infrastructure it runs on, to find root causes of inefficiency.