Research Engineer - AI Performance & Kernel Optimization
Core
Optimize performance of large-scale language model training and inference stacks by designing and implementing highly optimized kernels and eliminating bottlenecks across accelerator platforms.
Role type
Senior IC research engineer (AI performance & kernel optimization)
Builds
Optimized kernels and distributed training/inference systems for frontier-scale AI models
Domain
AI infrastructure, high-performance computing, GPU/accelerator systems
Deliverable
production ML models
Required skills
GPU kernel development (PTX, CUDA, HIP, Triton), low-level performance tuning, distributed training parallelism, memory hierarchy optimization, profiling and debugging, hardware-software interaction reasoning
Preferred skills
Non-NVIDIA hardware experience (AMD MI300x/MI355x, AWS Trainium, Google TPU), HPC background, compiler or numerical simulation experience
Technologies
CUDA, HIP, Triton, PTX, AMD MI300x, MI355x, AWS Trainium, Google TPU
Responsibilities
Develop and optimize GPU kernels for large-scale ML workloads, profile and eliminate bottlenecks in memory movement and communication, optimize distributed training for large MoE models, collaborate with research teams to translate system improvements into model gains
Seniority
Senior, hands-on IC