CareerPlanGet AI match score →

Research Engineer, Core ML

San Francisco💼 Full-time💰 $200,000–$200,000🗓 2026-07-10 → 2026-07-31

Core

Translate new RL algorithms, scheduling methods, and inference optimizations into production-grade systems that power Together's API, focusing on efficient inference and RL-driven training to improve latency, throughput, and model quality.

Role type

Staff-level Research Engineer (Core ML / Inference & RL Systems)

Builds

High-performance inference engines, RL/post-training pipelines, and serving systems for Together's API

Domain

Artificial Intelligence / Large Language Models / High-Performance Computing

Deliverable

production ML models

Required skills

Large-scale inference systems design, RL/post-training for LLMs, Transformer architecture design, distributed systems for ML, Python coding, GPU performance profiling, production system implementation

Preferred skills

Kernel backend development, speculative decoding, quantization, async RL rollouts, reward modeling, ablation study design

Technologies

SGLang, vLLM, FasterTransformer, TensorRT, ATLAS, GRPO, RLHF, RLAIF, DPO, FlashAttention, Hyena, FlexGen

Responsibilities

Design and prototype algorithms, architectures, and scheduling strategies for low-latency, high-throughput inference; Implement and maintain changes in high-performance inference engines including kernel backends and speculative decoding; Design and operate RL and post-training pipelines where inference cost dominates; Profile, debug, and optimize inference and post-training services under real production workloads; Set technical direction for cross-team efforts at the intersection of inference, RL, and post-training; Mentor other engineers and researchers on full-stack ML systems work

Seniority

Staff, technical leadership & mentorship

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗