CareerPlanSign in

Research Engineer - AI Performance & Kernel Optimization

San Francisco💼 Full-time🗓 2026-03-16 → 2026-09-26

Core

Optimize performance of large-scale language model training and inference stacks by designing and implementing highly optimized kernels and eliminating bottlenecks across accelerator platforms.

Role type

Senior IC research engineer (AI performance & kernel optimization)

Builds

Optimized kernels and distributed training/inference systems for frontier-scale AI models

Domain

AI infrastructure, high-performance computing, GPU/accelerator systems

Deliverable

production ML models

Required skills

GPU kernel development (PTX, CUDA, HIP, Triton), low-level performance tuning, distributed training parallelism, memory hierarchy optimization, profiling and debugging, hardware-software interaction reasoning

Preferred skills

Non-NVIDIA hardware experience (AMD MI300x/MI355x, AWS Trainium, Google TPU), HPC background, compiler or numerical simulation experience

Technologies

CUDA, HIP, Triton, PTX, AMD MI300x, MI355x, AWS Trainium, Google TPU

Responsibilities

Develop and optimize GPU kernels for large-scale ML workloads, profile and eliminate bottlenecks in memory movement and communication, optimize distributed training for large MoE models, collaborate with research teams to translate system improvements into model gains

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.