Staff Engineer, Datacenter Server Lifecycle
Core
Own the end-to-end operational journey of every machine in global datacenter facilities, from provisioning and deployment through steady-state operation, maintenance, repair, and decommissioning.
Role type
Staff Engineer, Datacenter Server Lifecycle
Builds
Automation for datacenters containing tens of thousands of servers; tooling to track machine health, configuration, and operational status.
Domain
AI infrastructure / Datacenter operations / Hardware lifecycle management
Deliverable
infrastructure
Required skills
Server hardware deployment and troubleshooting, Hardware lifecycle management, Programming (Python/Rust/Go/Java), Cloud infrastructure (Kubernetes, AWS/Azure/GCP), Trusted compute standards, Fleet management
Preferred skills
GPU/AI accelerator hardware experience, LinuxBoot/NixOS, OS distribution at scale, Capacity planning, Hardware security (TPM, attestation, secure boot)
Technologies
Kubernetes, AWS, Azure, GCP, LinuxBoot, NixOS, NVIDIA A100/H100, Google TPUs, AWS Trainium
Responsibilities
Build automation for large-scale datacenters; Define and own end-to-end system lifecycle strategy; Partner with Infrastructure Security to enforce trusted compute standards; Work with Networking team for end-to-end connectivity; Build tooling for fleet health tracking.
Seniority
Staff, hands-on IC with strategy ownership