Staff Software Engineer, GPU Inference
Core
Productionize and optimize GPU inference systems for large language models, focusing on disaggregated prefill/decode architectures and rack-scale AMD GPU infrastructure.
Role type
Staff Software Engineer (GPU Inference Systems)
Builds
GPU-accelerated inference stack (APIs, vLLM runtime, ROCm, rack-scale infrastructure)
Domain
AI Infrastructure / High-Performance Computing / GPU Systems
Deliverable
production ML models
Required skills
C++, Python, distributed systems debugging, GPU performance optimization, vLLM/SGLang/TensorRT-LLM, Linux/Kubernetes, numerical correctness validation, benchmarking
Preferred skills
AMD ROCm/HIP ecosystem, CUDA, open-source ML framework contributions, disaggregated prefill/decode architecture, KV-cache management, multi-GPU parallelism, quantization (BF16/FP8/INT8), kernel optimization
Technologies
vLLM, PyTorch, ROCm, HIP, Kubernetes, RDMA, BF16, FP8, INT8
Responsibilities
Design and deploy complete GPU prefill paths; establish operational readiness and automation for GPU fleets; define SLOs and improve fault isolation/recovery; profile and optimize throughput/latency/memory efficiency; tune scheduling and parallelism strategies; debug failures across application/runtime/hardware layers; build validation infrastructure for numerical correctness.
Seniority
Staff, hands-on IC with technical leadership