GPU Kernel Engineer – CUDA, Triton & Accelerator Performance
Core
Reviewing, debugging, and evaluating high-performance GPU and accelerator kernels for AI workloads to ensure numerical correctness, efficiency, and hardware suitability.
Role type
Senior IC GPU Kernel Engineer (CUDA/Triton/Accelerators)
Builds
Optimized GPU kernels, performance benchmarks, and technical evaluations for AI compute systems
Domain
AI Infrastructure / High-Performance Computing (HPC)
Deliverable
production ML models
Required skills
CUDA, Triton, NKI, Pallas, GPU performance optimization, kernel profiling, memory hierarchy optimization, numerical correctness verification, compilation debugging, operator fusion
Preferred skills
AWS Neuron, JAX, XLA, MLIR, compiler engineering, custom accelerator ecosystems (Trainium, TPU), cuBLAS, cuDNN
Technologies
CUDA, Triton, Nsight, NCU, roofline analysis, JAX, XLA, MLIR
Responsibilities
Reviewing kernel implementations for correctness and efficiency, debugging compilation and runtime issues, profiling and benchmarking kernel performance, translating kernels between frameworks, migrating kernels across hardware platforms, providing actionable technical feedback on optimization opportunities
Seniority
Senior, hands-on IC