Software Engineer, GPU Infrastructure- ChatGPT Engineering
Core
Design, build, and operate software systems that manage large-scale GPU clusters for ChatGPT inference, focusing on fleet health, capacity planning, and operational automation.
Role type
Senior IC GPU Infrastructure Engineer
Builds
Internal platforms, tooling, and AI-powered agents for GPU fleet operations
Domain
AI Infrastructure / High-Performance Computing / Distributed Systems
Deliverable
production ML models
Required skills
Go, Python, C++, Rust, Kubernetes, Linux, distributed systems design, observability, capacity planning, incident management
Preferred skills
GPU cluster operations, SRE, platform engineering, automation of operational workflows
Technologies
Kubernetes, Linux, cloud infrastructure, networking, observability tooling
Responsibilities
Design and build software for large-scale GPU infrastructure; develop internal platforms and AI agents to automate fleet operations; improve observability, reliability, and efficiency across thousands of GPUs; develop systems for capacity planning, scheduling, and incident response; identify infrastructure bottlenecks and implement solutions; partner with research and systems teams to improve the compute platform.
Seniority
Senior, hands-on IC