CareerPlanSign in

Software Engineer, Compute Foundations

San Francisco💼 Full-time🗓 2026-09-16 → 2026-09-25

Core

Build distributed systems that provision, configure, and manage OpenAI's GPU compute infrastructure across sites and data centers.

Role type

Senior IC distributed systems engineer (infrastructure control plane)

Builds

Kubernetes-based control planes, controllers, services, and APIs for machine lifecycle management

Domain

Cloud infrastructure, GPU compute, bare-metal systems

Deliverable

production ML models | infrastructure

Required skills

distributed systems design, Kubernetes APIs and reconciliation, bare-metal lifecycle management (PXE, BMC, firmware, Linux), API design, concurrency and failure handling, system reliability diagnostics

Preferred skills

multi-site control plane experience, GPU/HPC infrastructure knowledge, multi-provider hardware integration

Technologies

Kubernetes, Linux, bare-metal systems, network boot protocols

Responsibilities

Design and operate Kubernetes controllers for infrastructure coordination; Define APIs and resource models for lifecycle operations; Build provisioning services for hardware configuration and OS deployment; Develop lifecycle management for discovery, allocation, upgrades, and recovery; Design reliable reconciliation mechanisms for concurrent changes; Optimize control-plane throughput and API latency; Integrate new sites and GPU hardware generations into the platform

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.