•Run our cloud platform - design, build, and operate our GCP environments (production and sandbox) as infrastructure as code.
•Own Kubernetes - operate our GKE clusters, deploy pipelines, autoscale, and rollout strategy for the Rails monolith, Sidekiq workers, and supporting services.
•Keep data safe and fast - operate PostgreSQL (AlloyDB) and Redis: replication, read replicas, backups, point-in-time recovery, upgrades, and query and connection tuning.
•Make failure boring - build and test disaster recovery and failover. Define SLOs, alerts, and runbooks, and lead the response when an incident occurs.
•Give engineers visibility - own metrics, logs, and traces so that each team finds the cause of a problem in minutes.
•Ship fast and safe - improve CI/CD so that engineers deploy many times a day with zero downtime.
•Secure the platform - own network design, IAM, secrets management, and access control. Keep the platform compliant with PCI-DSS and regulator audits.
•Control cost - measure and reduce cloud spend without loss of reliability.
WHAT YOU WILL BRING
•Production infrastructure experience - 2-6 years of operation of production systems with real traffic, preferably in payments or fintech.
•Strong GCP skills - hands-on experience with GKE, VPC networks, IAM, Cloud SQL or AlloyDB, and Cloud Load Balancing. Deep AWS experience also counts.
•Kubernetes depth - run clusters in production and know Helm, autoscalers, resource limits, and safe rollouts.
•Database operations - operate PostgreSQL at scale: replication, failover, backups, vacuum, and index and query tuning.
•Infrastructure as code - build and maintain infrastructure with Terraform or an equivalent tool, and you review it like application code.
•Scripting and automation - write clean, tested scripts and tools in Bash, Python, Go, or Ruby.
•CI/CD expertise - build pipelines for trunk-based development with frequent, safe deploys.
•Security awareness - protect financial data and customer PII by default: least-privilege access, encryption, and secrets management.
•Ownership mentality - own systems from design to on-call, and you fix the root cause, not only the symptom.
•Cross-functional collaboration - work closely with product engineers, security, and compliance to deliver changes.
GOOD TO HAVE EXPERIENCE WITH
•PCI-DSS, SOC 2, or audits from a financial regulator
•Observability stacks such as OpenTelemetry, Prometheus, or Grafana
•Redis and Terraform
•Zero-trust access tools such as Teleport
•Rails and Sidekiq in production
•Disaster recovery drills and chaos tests