Research Engineer, Post-Training Inference
Core
Build and optimize inference engines and services for customizing open-source foundation models to downstream applications, enabling a seamless path from post-training to production serving.
Role type
Senior Research Engineer (Post-Training Inference)
Builds
Fine-tuning, Reinforcement Learning, and Evaluation services for open-source AI models
Domain
Artificial Intelligence / Machine Learning Systems / Large Language Models
Deliverable
production ML models
Required skills
Python, Go, modern inference engines (SGLang, vLLM, TensorRT-LLM), LLM fine-tuning methods, software engineering
Preferred skills
low-precision model serving (FP4/FP8), Multi-LoRA, RL training optimization, CUDA/Triton/CuTE kernel development, Kubernetes cluster management, open-source contributions
Technologies
SGLang, vLLM, TensorRT-LLM, CUDA, Triton, CuTE, Kubernetes, Python, Go
Responsibilities
Design and build systems for customizing open-source models; Build integrations between Model Shaping and Inference platforms; Add features to inference engines for large-scale post-training experiments; Ensure service stability and robustness via on-call rotation
Seniority
Senior, hands-on IC