AI Researcher, Core ML (Turbo)
Core
Design and operate high-performance inference and RL/post-training engines to make LLMs faster, cheaper, and more capable at production scale.
Role type
Staff-level AI Researcher (Core ML / Inference & RL Systems)
Builds
Production inference stacks (SGLang, vLLM-style), RL training pipelines (GRPO, RLHF), and efficient model serving systems.
Domain
Artificial Intelligence, Large Language Models, High-Performance Computing
Deliverable
production ML models
Required skills
Large-scale inference systems, RL/post-training for LLMs, Transformer architecture design, Distributed systems for ML, Python, GPU performance profiling, End-to-end project ownership
Preferred skills
Kernel backend implementation, Speculative decoding, Quantization, Async RL rollouts, Ablation study design
Technologies
SGLang, vLLM, FasterTransformer, TensorRT, ATLAS, GRPO, RLHF, RLAIF, DPO
Responsibilities
Design and prototype algorithms for low-latency, high-throughput inference; Implement and maintain changes in high-performance inference engines; Design and operate RL and post-training pipelines; Profile, debug, and optimize inference and post-training services under real production workloads; Set technical direction for cross-team efforts at the intersection of inference, RL, and post-training.
Seniority
Staff, technical leadership & mentorship