CareerPlanGet AI match score →

Software Engineer, Inference – AMD GPU Enablement

San Francisco💼 Full-time🗓 2025-10-08 → 2026-07-31

Core

Scaling and optimizing OpenAI's inference infrastructure across emerging GPU platforms, specifically focusing on AMD accelerators to ensure large models run smoothly.

Role type

Senior IC machine-learning infrastructure engineer (GPU inference)

Builds

High-performance distributed inference systems on AMD hardware

Domain

AI/ML inference, GPU computing, distributed systems

Deliverable

production ML models

Required skills

GPU kernel development (HIP, CUDA, Triton), distributed inference systems, communication libraries (NCCL, RCCL), system-level debugging and optimization, memory and compute profiling

Preferred skills

Open-source contributions (RCCL, Triton, vLLM), GPU performance profiling tools (Nsight, rocprof), non-NVIDIA GPU deployment experience, model/tensor parallelism knowledge

Technologies

HIP, CUDA, Triton, vLLM, RCCL, NCCL, ROCm

Responsibilities

Own bring-up, correctness, and performance of the inference stack on AMD hardware; Integrate internal model-serving infrastructure into GPU-backed systems; Debug and optimize distributed inference workloads; Validate correctness and scalability on large GPU clusters; Design and optimize high-performance GPU kernels; Build and tune collective communication libraries for parallelization

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗