CareerPlanSign in

Software Engineer - GPU Kernels

San Francisco💼 Full-time🗓 2025-07-17 → 2026-09-26

Core

Design and implement high-performance GPU kernels for key ML operations (matrix multiplications, attention mechanisms, MoE routing) to optimize computation for state-of-the-art machine learning models.

Role type

Senior IC GPU Kernel Engineer

Builds

High-performance inference stack and GPU libraries for AI workloads

Domain

AI Infrastructure / GPU Computing

Deliverable

production ML models

Required skills

CUDA C++ API, GPU architecture (memory hierarchy, thread/block/grid), performance profiling (Nsight Systems/Compute), memory coalescing, warp-level programming, tensor core acceleration, compute/memory overlap, C++

Preferred skills

Transformer model optimization (Flash Attention), GPU kernel libraries (Cutlass, Triton, Thrust, CUB), GEMM tuning, distributed/multi-GPU compute, open-source GPU contributions, research publications

Technologies

CUDA, PTX, Nsight Systems, Nsight Compute, Torch Profiler, C++

Responsibilities

Design and implement high-performance GPU kernels for ML operations; Write and optimize code using CUDA and PTX assembly; Apply advanced performance optimization methods; Implement cutting-edge features like quantization and sparsity; Identify and resolve performance bottlenecks using profiling tools; Contribute to internal and open-source GPU libraries

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.