CareerPlanGet AI match score →

Member of Technical Staff - ML Training Systems

New York💼 Full-time🗓 2026-02-25 → 2026-08-01

Core

Building the infrastructure layer for training production machine learning models, specifically optimizing for large language models.

Role type

Senior IC ML Training Systems Engineer

Builds

High-performance training infrastructure for language models

Domain

Cloud infrastructure + Machine Learning

Deliverable

production ML models

Required skills

high-performance code, PyTorch, Hugging Face, ML training optimization, data loading optimization, communication/computation overlap, off-policy rollouts

Preferred skills

Linux kernel, file systems, containers

Technologies

torch, Huggingface, verl, slime

Responsibilities

Write high-quality, high-performance code for ML training; Optimize data loading and communication patterns; Participate in on-call rotation for production incidents

Seniority

Senior, hands-on IC

Rewrite
## About Us AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn, Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. ## The Role We are looking for strong engineers with experience training production machine learning models. If you are interested in contributing to open-source projects and evolving Modal's infrastructure to train the next generation of language models, we'd love to hear from you! ## Requirements - 5+ years of experience writing high-quality, high-performance code. - Experience working with torch and high-level training frameworks (Huggingface, verl, slime) - Experience with ML training optimization (tell us a story about eliminating data loading bottlenecks, overlapping communications with compute, rewriting a trainer to handle off-policy rollouts, etc.) - Ability to work in-person, in our NYC or San Francisco office. - Ability to participate in on-call rotation and respond to production incidents. ## Nice to Have - Familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc). ## Benefits - Working with leading-edge AI infrastructure. - Collaborating with top-tier engineers and researchers. - Contributing to open-source projects. - Opportunities for growth and impact.
Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗