CareerPlanSign in

Post-Training Research Engineer

San Francisco💼 Full-time🗓 2026-03-23 → 2026-09-26

Core

Build in-house tooling to support post-training of custom AI models (RL, SFT, distillation) for customers, focusing on efficiency and quality at scale.

Role type

Senior IC research engineer (post-training infrastructure)

Builds

Scalable training tooling for transformer models using distributed systems and GPU kernels

Domain

AI infrastructure / Distributed ML systems

Deliverable

production ML models

Required skills

PyTorch, transformer training parallelism strategies, distributed GPU program profiling, roofline analysis, HPC/distributed computing platforms, cluster networking, operating systems fundamentals

Preferred skills

Kubernetes, cgroups, storage systems, networking topologies, Jax/TensorFlow, Infiniband/RoCE/GPUDirect

Technologies

PyTorch, Kubernetes, Slurm, Ray, Dask, Infiniband, RoCE, GPUDirect

Responsibilities

Develop tooling for training diverse model architectures with various techniques; optimize distributed GPU programs; perform roofline analysis on training setups; collaborate with researchers to derive specifications

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.