CareerPlanSign in

GPU Kernel Engineer

San Francisco, CA💼 Full-time🗓 2026-09-25 → 2026-09-26

Core

Design and optimize custom GPU kernels to power next-generation large-scale AI systems, specifically focusing on large-scale LLM training and inference.

Role type

Senior IC GPU Kernel Engineer

Builds

Optimized GPU kernels integrated into high-level ML frameworks (PyTorch, JAX) for frontier AI models and real-time applications

Domain

AI Infrastructure / High-Performance Computing / GPU Programming

Deliverable

production ML models

Required skills

C++, CUDA, ROCm, Triton, JAX Pallas, PTX, GPU memory models, performance profiling, low-level GPU execution, ML framework integration

Preferred skills

AMD GPU optimization, JAX FFI, model serving frameworks (vLLM, TensorRT), TPU/XLA programming, open-source contributions

Technologies

C++, Python, CUDA, ROCm, Triton, JAX, PyTorch, PTX, vLLM, TensorRT

Responsibilities

Design, implement, and optimize custom GPU kernels; Profile and optimize end-to-end performance of ML operations; Integrate low-level GPU kernels into frameworks; Develop performance models and identify bottlenecks; Collaborate with ML researchers and distributed systems engineers; Work with hardware vendors on architecture capabilities; Contribute to tooling, documentation, and testing frameworks

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.