Software Engineer, Inference – AMD GPU Enablement
Core
Scaling and optimizing OpenAI's inference infrastructure across emerging GPU platforms, specifically focusing on AMD accelerators to ensure large models run smoothly.
Role type
Senior IC machine-learning infrastructure engineer (GPU inference)
Builds
High-performance distributed inference systems on AMD hardware
Domain
AI/ML inference, GPU computing, distributed systems
Deliverable
production ML models
Required skills
GPU kernel development (HIP, CUDA, Triton), distributed inference systems, communication libraries (NCCL, RCCL), system-level debugging and optimization, memory and compute profiling
Preferred skills
Open-source contributions (RCCL, Triton, vLLM), GPU performance profiling tools (Nsight, rocprof), non-NVIDIA GPU deployment experience, model/tensor parallelism knowledge
Technologies
HIP, CUDA, Triton, vLLM, RCCL, NCCL, ROCm
Responsibilities
Own bring-up, correctness, and performance of the inference stack on AMD hardware; Integrate internal model-serving infrastructure into GPU-backed systems; Debug and optimize distributed inference workloads; Validate correctness and scalability on large GPU clusters; Design and optimize high-performance GPU kernels; Build and tune collective communication libraries for parallelization
Seniority
Senior, hands-on IC