SDE IV - GPU Engineer
Core
Architecting and optimizing GPU inference runtimes for large-scale diffusion and transformer models to enable hyper-realistic, personal shopping experiences.
Role type
Senior IC GPU Systems Engineer (Inference Stack)
Builds
High-performance inference runtimes, kernel dispatchers, memory planners, and distributed training/inference systems for Stable Diffusion and multimodal transformers.
Domain
AI Commerce / High-Performance Computing / GPU Infrastructure
Deliverable
production ML models
Required skills
CUDA, Triton, C++, GPU scheduling, tensor cores, distributed inference systems, NCCL, NVLink, PCIe, interconnects, profiling automation, technical leadership
Preferred skills
compiler-aided optimization (TVM, XLA, MLIR), Stable Diffusion inference tuning, heterogeneous compute backends (AMD ROCm, TPU, ASICs), hardware–software co-design
Technologies
CUDA, Triton, C++, NCCL, NVLink, PCIe, TVM, XLA, MLIR, ROCm, TPU
Responsibilities
Architect high-performance inference runtimes and kernel dispatchers; investigate cross-GPU performance bottlenecks; drive multi-GPU parallelism strategies; establish GPU optimization standards and tooling; collaborate with research on novel architectures; mentor engineers in low-level optimization; partner with hardware vendors to maximize cluster utilization.
Seniority
Senior, hands-on IC with technical leadership