Staff Engineer, Datacenter Server Lifecycle
Core
Own the end-to-end operational journey of datacenter servers from provisioning to decommissioning, with a focus on automation, trusted compute standards, and fleet health tracking.
Role type
Staff Engineer, Datacenter Server Lifecycle
Builds
Automated server lifecycle processes, tooling for fleet health tracking, and trusted compute standards for AI hardware.
Domain
Cloud Infrastructure / Datacenter Operations / Hardware Security
Deliverable
infrastructure
Required skills
Server hardware lifecycle management, automation development, programming (Python/Rust/Go/Java), cloud infrastructure (Kubernetes, IAC, AWS, GCP), hardware troubleshooting, cross-functional collaboration
Preferred skills
GPU/AI accelerator hardware experience, provisioning tooling (coreboot/LinuxBoot/u-root), fleet management platform development, OS distribution at scale, capacity planning, trusted compute concepts (TPM, attestation, secure boot)
Technologies
Kubernetes, Infrastructure as Code, AWS, GCP, Python, Rust, Go, Java, coreboot, LinuxBoot, u-root, NVIDIA A100/H100, AMD MI300, Google TPUs, AWS Trainium
Responsibilities
Lead build-out of automation for tens of thousands of servers; Define and own end-to-end server lifecycle strategy; Partner with Infrastructure Security to enforce trusted compute standards; Work with Networking team for end-to-end connectivity; Build and maintain tooling to track machine health and configuration. (via careerplan.io/jobs/5139038008-staff-engineer-datacenter-server-lifecycle-at-anthropic)
Seniority
Staff, hands-on IC with strategic scope
