Performance Engineer, GPU
Core
Architect and implement foundational systems to maximize GPU utilization and performance at unprecedented scale for large language models.
Role type
Senior IC GPU Performance Engineer
Builds
Production ML systems powering Claude and frontier language models
Domain
AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
GPU programming and optimization at scale, CUDA, Triton, CUTLASS, Flash Attention, tensor core optimization, PyTorch/JAX internals, kernel fusion, memory bandwidth optimization, profiling with Nsight, NCCL, NVLink, collective communication, model parallelism, INT8/FP8 quantization, mixed-precision techniques, large-scale training infrastructure, fault tolerance, cluster orchestration
Preferred skills
Co-designing attention mechanisms for next-gen hardware, developing custom kernels for emerging quantization formats, designing distributed communication strategies, optimizing end-to-end training/inference pipelines, building performance modeling frameworks, implementing kernel fusion strategies, creating resilient planet-scale distributed training systems, partnering with hardware vendors
Technologies
CUDA, Triton, CUTLASS, Flash Attention, PyTorch, JAX, torch.compile, XLA, Nsight, NCCL, NVLink
Responsibilities
Architect and implement foundational systems for GPU performance, develop custom kernels for quantization and mixed-precision, design distributed communication strategies for multi-node clusters, optimize end-to-end training and inference pipelines, build performance modeling frameworks, implement kernel fusion strategies, create resilient distributed training infrastructure, profile and eliminate bottlenecks in production serving, partner with hardware vendors
Seniority
Senior, hands-on IC