Sr. Engineering Manager, AI Runtime
Core
Lead the team owning the product experience and foundational infrastructure for Databricks' AI Runtime (AIR), enabling enterprises to train and fine-tune deep learning and LLM models with on-demand GPUs.
Role type
Senior Engineering Manager, AI Runtime Infrastructure
Builds
Managed GPU training infrastructure for deep learning and LLM model development
Domain
Cloud AI Infrastructure / Distributed Deep Learning
Deliverable
production ML models
Required skills
Engineering management, distributed training frameworks (PyTorch, DeepSpeed, Megatron-LM), GPU performance optimization, training resilience patterns, cross-functional leadership, product roadmap definition
Preferred skills
Experience with FSDP, tensor/pipeline parallelism, NCCL, interconnect topologies, memory optimization
Technologies
PyTorch, DeepSpeed, Composer, Megatron-LM, NCCL
Responsibilities
Lead and mentor a high-performing engineering team for Custom Training product and infrastructure; Define and own the product and technical roadmap for AIR; Collaborate with product, research, and platform teams to drive end-to-end delivery; Drive architectural decisions for managed GPU training at scale; Build observability and reliability practices for long-running training jobs; Partner with recruiting to hire top-tier engineering talent
Seniority
Senior, hands-on IC with management responsibilities