Senior Site Reliability Engineer I
Core
Architect, operate, and scale high-performance NGINX ingress fleets and Kubernetes infrastructure for Braze's Ruby on Rails monolith and Go API services, ensuring maximum uptime for 3.3 billion monthly active users.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
High-throughput API ingestion layers and scalable routing systems for a global customer engagement platform
Domain
Internet-scale distributed systems, API infrastructure, Cloud-native technologies
Deliverable
production ML models | product features | infrastructure
Required skills
NGINX configuration and operation, Kubernetes administration, Linux internals, Ruby or Go programming, Infrastructure as Code (Terraform/Ansible), Systems design
Preferred skills
Redis, Kafka, Postgres, MongoDB, Prometheus, Grafana, Datadog, AWS/GCP/Azure
Responsibilities
Configure and tune NGINX routing and ingress controllers, manage automated scaling routines using RED metrics and HPA, establish SLIs/SLOs and error budgets, conduct root-cause analysis and blameless retrospectives, perform capacity planning and bottleneck profiling
Seniority
Senior, hands-on IC
