Software Engineer, Compute Foundations
Core
Build distributed systems that provision, configure, and manage OpenAI's GPU compute infrastructure across sites and data centers.
Role type
Senior IC distributed systems engineer (infrastructure control plane)
Builds
Kubernetes-based control planes, controllers, services, and APIs for machine lifecycle management
Domain
Cloud infrastructure, GPU compute, bare-metal systems
Deliverable
production ML models | infrastructure
Required skills
distributed systems design, Kubernetes APIs and reconciliation, bare-metal lifecycle management (PXE, BMC, firmware, Linux), API design, concurrency and failure handling, system reliability diagnostics
Preferred skills
multi-site control plane experience, GPU/HPC infrastructure knowledge, multi-provider hardware integration
Technologies
Kubernetes, Linux, bare-metal systems, network boot protocols
Responsibilities
Design and operate Kubernetes controllers for infrastructure coordination; Define APIs and resource models for lifecycle operations; Build provisioning services for hardware configuration and OS deployment; Develop lifecycle management for discovery, allocation, upgrades, and recovery; Design reliable reconciliation mechanisms for concurrent changes; Optimize control-plane throughput and API latency; Integrate new sites and GPU hardware generations into the platform