Lead Site Reliability Engineer (SRE)
Core
Lead a small SRE team to define long-term infrastructure vision, ensuring systems remain reliable, scalable, and secure while handling massive traffic influxes.
Role type
Lead Site Reliability Engineer (SRE)
Builds
Distributed databases, high-throughput monitoring infrastructure, CI/CD pipelines, and automated cluster management tools.
Domain
Cloud infrastructure, distributed systems, high-traffic production environments
Deliverable
production ML models | product features | infrastructure
Required skills
System architecture design, low-level language programming (Rust or Go), Kubernetes management, cloud platform proficiency (AWS/GCP), monitoring and alerting tool expertise, CI/CD pipeline development, high-traffic scaling strategies
Preferred skills
Experience with ScyllaDB, Redpanda, GCP partners, automated Kubernetes Pod Autoscaler implementation
Technologies
ScyllaDB, Kubernetes, AWS, GCP, Rust, Go, Redpanda
Responsibilities
Define and lead long-term infrastructure roadmap, guide SRE team execution, design and maintain large-scale distributed databases, automate manual operational steps, partner with backend engineers on code production, build new tools to accelerate company growth
Seniority
Senior, hands-on IC with leadership responsibilities