Research Engineer, Core ML
Core
Translate new RL algorithms, scheduling methods, and inference optimizations into production-grade systems that power Together's API, focusing on efficient inference and RL-driven training to improve latency, throughput, and model quality.
Role type
Staff-level Research Engineer (Core ML / Inference & RL Systems)
Builds
High-performance inference engines, RL/post-training pipelines, and serving systems for Together's API
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
Large-scale inference systems design, RL/post-training for LLMs, Transformer architecture design, distributed systems for ML, Python coding, GPU performance profiling, production system implementation
Preferred skills
Kernel backend development, speculative decoding, quantization, async RL rollouts, reward modeling, ablation study design
Technologies
SGLang, vLLM, FasterTransformer, TensorRT, ATLAS, GRPO, RLHF, RLAIF, DPO, FlashAttention, Hyena, FlexGen
Responsibilities
Design and prototype algorithms, architectures, and scheduling strategies for low-latency, high-throughput inference; Implement and maintain changes in high-performance inference engines including kernel backends and speculative decoding; Design and operate RL and post-training pipelines where inference cost dominates; Profile, debug, and optimize inference and post-training services under real production workloads; Set technical direction for cross-team efforts at the intersection of inference, RL, and post-training; Mentor other engineers and researchers on full-stack ML systems work
Seniority
Staff, technical leadership & mentorship