Site Reliability Engineer
Core
Designated Responsible Individual (DRI) monitoring service health, managing on-call rotations, and implementing solutions for performance and functionality issues in large-scale distributed systems.
Role type
Senior Site Reliability Engineer (Infrastructure & Distributed Systems)
Builds
Production and deployment automation for complex product features; mitigations for Live Site service issues.
Domain
Cloud infrastructure, distributed systems, GPU/InfiniBand hardware support
Deliverable
production ML models | infrastructure
Required skills
On-call incident management, data collection and classification for system health, automation development, physical infrastructure management, large-scale cloud/distributed systems experience, project lifecycle ownership
Preferred skills
People management, broad influence and communication, end-to-end project management
Technologies
Cloud platforms, distributed systems, GPUs, InfiniBand
Responsibilities
Monitor service for degradation and downtime; collect and analyze system metrics; develop automation for production and deployment; implement solutions for complex performance issues; manage physical infrastructure supporting GPUs and InfiniBand.
Seniority
Senior, hands-on IC with people management experience