ML Research Engineer, ML Systems
Core
Building and optimizing Scale's internal distributed framework (RLXF) for large language model training, inference, and data quality evaluation.
Role type
ML Research Engineer, ML Systems
Builds
Distributed training and inference framework for LLMs
Domain
Artificial Intelligence / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
CUDA, PyTorch, Transformers, Flash Attention, multi-node LLM training, distributed ML systems, software engineering
Preferred skills
post-training methods, instruction tuning, RLHF, tool use, reasoning, agents, multimodal models
Technologies
CUDA, PyTorch, Transformers, Flash Attention
Responsibilities
Build, profile, and optimize training and inference framework; Collaborate with ML teams to accelerate research and development; Research and integrate state-of-the-art technologies to optimize ML systems
Seniority
Individual Contributor