Senior Site Reliability Engineer (Hardware Automation)
Core
Building internal platforms and tooling to automate data center infrastructure operations, ensuring fault-tolerance and scale for AI cloud services.
Role type
Senior Site Reliability Engineer (Hardware Automation)
Builds
Internal automation platforms and tooling for hardware infrastructure teams
Domain
Data center operations / Hardware infrastructure / AI cloud platform
Deliverable
infrastructure
Required skills
Linux systems administration, Python scripting, Bash scripting, complex system troubleshooting, analytical problem-solving
Preferred skills
Backend development, high-load distributed systems design
Technologies
Linux, Python, Bash
Responsibilities
Ensure fault-tolerance, scale and uninterrupted operations for services; Use cutting-edge technology to solve infrastructure problems; Implement and improve CI/CD processes
Seniority
Senior, hands-on IC
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.