•Lead, hire, onboard, and develop a distributed engineering team working asynchronously.
•Set priorities with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams, and help the team deliver observability services iteratively.
•Own the reliability, scalability, and cost of the team's metrics, logging, alerting, and capacity planning platforms.
•Reduce noisy or missing alerts and telemetry gaps, and use SLOs, error budgets, and self-service instrumentation to help engineers maintain the health of their services.
•Guide technical decisions about time-series storage, high-cardinality metrics, log pipelines, and distributed tracing.
•Participate in the Incident Manager On Call (IMOC) rotation, coordinating the response to high-severity incidents affecting GitLab.com .
•Keep the team's on-call rotation sustainable through coverage across time zones, useful runbooks, better alerts, and follow-through on post-incident actions.
•Use AI tools and agents to support engineering workflows and incident triage, reviewing their output while engineers retain responsibility for decisions.
What you’ll bring
•Experience leading an observability, platform engineering, or site reliability engineering team operating at scale, including supporting people in a distributed, asynchronous environment.
•Technical knowledge of metrics systems such as Prometheus and long-term storage, logging platforms such as Elasticsearch or cloud-native services, and alerting design.
•Experience using SLOs, error budgets, and capacity forecasts to make reliability and investment decisions.
•Experience operating a large software-as-a-service platform and investigating production issues such as telemetry gaps, ingestion limits, or noisy and missing alerts.
•Experience participating in and improving production on-call rotations, including incident coordination and balancing operational load with project work.
•The ability to explain technical tradeoffs to engineering partners and other stakeholders.
•Experience using AI tools or agents in engineering or management work; you can describe how you would apply them to operational problems such as incident triage.
We welcome different paths into this role, whether through practical experience, formal study, or transferable skills. If the work interests you, please apply even if you don't meet every qualification.
About the team
We're part of Production Engineering within Infrastructure Platforms and work asynchronously across regions. Our tools include Tamland, the capacity forecasting tool. We bring lessons from operating GitLab's production systems back into the platforms we build.
How GitLab Supports Full-Time Employees
•Benefits to support your health, finances, and well-being
•Flexible Paid Time Off
•Team Member Resource Groups
•Equity Compensation & Employee Stock Purchase Plan
•Growth and Development Fund
•Parental Leave
Please note that we welcome interest from candidates with varying levels of experience; many successful candidates do not meet every single requirement. Additionally, studies have shown that people from underrepresented groups are less likely to apply to a job unless they meet every single qualification. If you're excited about this role, please apply and allow our recruiters to assess your application.
Country Hiring Guidelines: GitLab hires new team members in countries around the world. All of our roles are remote, however some roles may carry specific location-based eligibility requirements. Our Talent Acquisition team can help answer any questions about location after starting the recruiting process.
Privacy Policy: Please review our Recruitment Privacy Policy. Your privacy is important to us.
GitLab is proud to be an equal opportunity workplace and is an affirmative action employer. GitLab’s policies and practices relating to recruitment, employment, career development and advancement, promotion, and retirement are based solely on merit, regardless of race, color, religion, ancestry, sex (including pregnancy, lactation, sexual orientation, gender identity, or gender expression), national origin, age, citizenship, marital status, mental or physical disability, genetic information (including family medical history), discharge status from the military, protected veteran status (which includes disabled veterans, recently separated veterans, active duty wartime or campaign badge veterans, and Armed Forces service medal veterans), or any other basis protected by law. GitLab will not tolerate discrimination or harassment based on any of these characteristics. See also GitLab’s EEO Policy and EEO is the Law . If you have a disability or special need that requires accommodation , please let us know during the recruiting process .