Senior Distributed ML Engineer
Core
Senior distributed ML engineer optimizing training/inference for very large models in AI safety research.
Role type
Senior IC distributed ML engineer (research infrastructure)
Builds
Distributed training/inference frameworks and tooling for large-scale AI safety models
Domain
AI Safety / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
distributed ML training frameworks, GPU profiling, cloud platforms, workload managers, containerization, deep learning research
Preferred skills
advanced degree in ML/distributed systems, experience with Megatron/DeepSpeed/FSDP/vLLM/verl
Technologies
Megatron, DeepSpeed, HuggingFace Accelerate, FSDP, vLLM, verl, Ray, SLURM, PyTorch profiler, PyProf, NVIDIA Nsight, Docker, Kubernetes, gRPC
Responsibilities
Collaborate with researchers to accelerate model training and inference; investigate performance bottlenecks and debug issues; develop tools for distributed computing orchestration; establish best practices for large-scale ML workflows
Seniority
Senior, hands-on IC