GPU Kernel Engineer
Core
Design and optimize custom GPU kernels to power next-generation large-scale AI systems, specifically focusing on large-scale LLM training and inference.
Role type
Senior IC GPU Kernel Engineer
Builds
Optimized GPU kernels integrated into high-level ML frameworks (PyTorch, JAX) for frontier AI models and real-time applications
Domain
AI Infrastructure / High-Performance Computing / GPU Programming
Deliverable
production ML models
Required skills
C++, CUDA, ROCm, Triton, JAX Pallas, PTX, GPU memory models, performance profiling, low-level GPU execution, ML framework integration
Preferred skills
AMD GPU optimization, JAX FFI, model serving frameworks (vLLM, TensorRT), TPU/XLA programming, open-source contributions
Technologies
C++, Python, CUDA, ROCm, Triton, JAX, PyTorch, PTX, vLLM, TensorRT
Responsibilities
Design, implement, and optimize custom GPU kernels; Profile and optimize end-to-end performance of ML operations; Integrate low-level GPU kernels into frameworks; Develop performance models and identify bottlenecks; Collaborate with ML researchers and distributed systems engineers; Work with hardware vendors on architecture capabilities; Contribute to tooling, documentation, and testing frameworks
Seniority
Senior, hands-on IC
