Post-Training Research Engineer
Core
Build in-house tooling to support post-training of custom AI models (RL, SFT, distillation) for customers, focusing on efficiency and quality at scale.
Role type
Senior IC research engineer (post-training infrastructure)
Builds
Scalable training tooling for transformer models using distributed systems and GPU kernels
Domain
AI infrastructure / Distributed ML systems
Deliverable
production ML models
Required skills
PyTorch, transformer training parallelism strategies, distributed GPU program profiling, roofline analysis, HPC/distributed computing platforms, cluster networking, operating systems fundamentals
Preferred skills
Kubernetes, cgroups, storage systems, networking topologies, Jax/TensorFlow, Infiniband/RoCE/GPUDirect
Technologies
PyTorch, Kubernetes, Slurm, Ray, Dask, Infiniband, RoCE, GPUDirect
Responsibilities
Develop tooling for training diverse model architectures with various techniques; optimize distributed GPU programs; perform roofline analysis on training setups; collaborate with researchers to derive specifications
Seniority
Senior, hands-on IC