Research Engineer, Infrastructure, Kernels
Core
Design, optimize, and maintain compute foundations and high-performance ML kernels for large-scale language model training.
Role type
Senior IC infrastructure research engineer (kernels/systems)
Builds
Custom ML kernels, compute primitives, and performance benchmarks for internal model training
Domain
AI infrastructure / GPU computing / Deep Learning Systems
Deliverable
production ML models
Required skills
CUDA programming, CuTe, Triton, GPU architecture optimization, deep learning frameworks (PyTorch, JAX), compute profiling, low-precision arithmetic, distributed compute stacks
Preferred skills
Large-scale LLM training experience, tensor/pipeline parallelism, low-precision formats (FP8, INT8), compiler stacks (XLA, TVM), open-source GPU/ML contributions
Technologies
CUDA, CuTe, Triton, PyTorch, JAX, XLA, TVM
Responsibilities
Design and implement custom ML kernels for core LLM operations (attention, matrix multiplication, gating, normalization); Design compute primitives to reduce memory bandwidth bottlenecks; Collaborate with research teams to align kernel optimizations with model architecture; Develop and maintain a library of reusable kernels and performance benchmarks; Contribute to infrastructure stability and scalability; Document and share insights through internal talks, technical papers, or open-source contributions
Seniority
Senior, hands-on IC