Site Reliability Engineer
Core
Ensuring fault-tolerance, scale, and uninterrupted operations for data-center hardware and AI infrastructure systems.
Role type
Site Reliability Engineer (Hardware Infrastructure)
Builds
Data-center lifecycle systems including power, cooling, racks, servers, and network devices.
Domain
AI Cloud Infrastructure / Data Center Operations
Deliverable
infrastructure
Required skills
Linux systems administration, Python scripting, Bash scripting, complex system troubleshooting, analytical problem-solving
Preferred skills
backend development, high-load distributed systems design
Technologies
Linux, Python, Bash
Responsibilities
Monitor engineering equipment (power, cooling) and IT hardware (racks, servers, network devices), implement and improve CI/CD processes, track asset and hardware repairs.
Seniority
Individual Contributor
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.