CareerPlanGet AI match score →

Research Engineer, ML Systems (All Industry Levels)

Redwood City, CA💼 Full-time🗓 2024-11-20 → 2026-07-31

Core

Designing and optimizing high-performance ML training and inference systems for GPU clusters to serve LLMs at scale.

Role type

Research Engineer, ML Systems

Builds

GPU clusters, distributed RLHF stacks, multimodal model training/inference systems

Domain

Artificial Intelligence / Machine Learning Systems

Deliverable

production ML models

Required skills

PyTorch, distributed machine learning, reinforcement learning, transformers, Triton kernel development, CUDA programming, LLM training and distillation

Preferred skills

DeepSpeed, Megatron, vLLM, FlashAttention, Kubernetes, Docker, cloud orchestration, academic publications

Responsibilities

Write efficient Triton kernels for specific models and hardware, develop prefix-aware routing algorithms, train and distill LLMs, build distributed RLHF stacks, develop systems for multimodal model training and inference

Seniority

All Industry Levels (PhD or equivalent research experience required)

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗