Software Engineer - GPU Kernels
Core
Design and implement high-performance GPU kernels for key ML operations (matrix multiplications, attention mechanisms, MoE routing) to optimize computation for state-of-the-art machine learning models.
Role type
Senior IC GPU Kernel Engineer
Builds
High-performance inference stack and GPU libraries for AI workloads
Domain
AI Infrastructure / GPU Computing
Deliverable
production ML models
Required skills
CUDA C++ API, GPU architecture (memory hierarchy, thread/block/grid), performance profiling (Nsight Systems/Compute), memory coalescing, warp-level programming, tensor core acceleration, compute/memory overlap, C++
Preferred skills
Transformer model optimization (Flash Attention), GPU kernel libraries (Cutlass, Triton, Thrust, CUB), GEMM tuning, distributed/multi-GPU compute, open-source GPU contributions, research publications
Technologies
CUDA, PTX, Nsight Systems, Nsight Compute, Torch Profiler, C++
Responsibilities
Design and implement high-performance GPU kernels for ML operations; Write and optimize code using CUDA and PTX assembly; Apply advanced performance optimization methods; Implement cutting-edge features like quantization and sparsity; Identify and resolve performance bottlenecks using profiling tools; Contribute to internal and open-source GPU libraries
Seniority
Senior, hands-on IC
