Senior Site Reliability Engineer
Core
Ensure reliability, performance, and resilience of high-performance infrastructure and AI inference platforms serving developers and businesses.
Role type
Senior Site Reliability Engineer (distributed systems & GPU workloads)
Builds
Production-grade GPU infrastructure, serverless AI model deployment platforms, and scalable inference services
Domain
Cloud infrastructure, AI/ML workloads, distributed systems
Deliverable
production ML models | infrastructure
Required skills
distributed systems debugging, observability (metrics/logs/tracing), SRE principles (SLIs/SLOs/error budgets), Kubernetes, containers, IaC, automation scripting (Python/Go/PHP), incident management, capacity planning
Preferred skills
high-throughput/low-latency API operations, bare-metal infrastructure, GPU environments, AI/ML workloads, RabbitMQ, MySQL/Redis/ClickHouse, global traffic management, automated scaling/self-healing systems
Responsibilities
Own reliability/availability/performance of critical production services; define and evolve reliability practices (SLIs, SLOs, alerting); investigate complex production issues across distributed systems and GPU workloads; lead incident reviews and RCAs; reduce operational toil via automation and deployment safety improvements; collaborate on capacity planning and architectural scaling
Seniority
Senior, hands-on IC