CareerPlanGet AI match score →

Research Engineer, Post-Training Inference

San Francisco💼 Full-time💰 $200,000–$200,000🗓 2026-07-10 → 2026-07-31

Core

Build and optimize inference engines and services for customizing open-source foundation models to downstream applications, enabling a seamless path from post-training to production serving.

Role type

Senior Research Engineer (Post-Training Inference)

Builds

Fine-tuning, Reinforcement Learning, and Evaluation services for open-source AI models

Domain

Artificial Intelligence / Machine Learning Systems / Large Language Models

Deliverable

production ML models

Required skills

Python, Go, modern inference engines (SGLang, vLLM, TensorRT-LLM), LLM fine-tuning methods, software engineering

Preferred skills

low-precision model serving (FP4/FP8), Multi-LoRA, RL training optimization, CUDA/Triton/CuTE kernel development, Kubernetes cluster management, open-source contributions

Technologies

SGLang, vLLM, TensorRT-LLM, CUDA, Triton, CuTE, Kubernetes, Python, Go

Responsibilities

Design and build systems for customizing open-source models; Build integrations between Model Shaping and Inference platforms; Add features to inference engines for large-scale post-training experiments; Ensure service stability and robustness via on-call rotation

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗