CareerPlanSign in

Sr. Engineering Manager, AI Runtime

San Francisco, California💼 Full-time💰 $228,600–$228,600🗓 2026-07-06 → 2026-09-26

Core

Lead the team owning the product experience and foundational infrastructure for Databricks' AI Runtime (AIR), enabling enterprises to train and fine-tune deep learning and LLM models with on-demand GPUs.

Role type

Senior Engineering Manager, AI Runtime Infrastructure

Builds

Managed GPU training infrastructure for deep learning and LLM model development

Domain

Cloud AI Infrastructure / Distributed Deep Learning

Deliverable

production ML models

Required skills

Engineering management, distributed training frameworks (PyTorch, DeepSpeed, Megatron-LM), GPU performance optimization, training resilience patterns, cross-functional leadership, product roadmap definition

Preferred skills

Experience with FSDP, tensor/pipeline parallelism, NCCL, interconnect topologies, memory optimization

Technologies

PyTorch, DeepSpeed, Composer, Megatron-LM, NCCL

Responsibilities

Lead and mentor a high-performing engineering team for Custom Training product and infrastructure; Define and own the product and technical roadmap for AIR; Collaborate with product, research, and platform teams to drive end-to-end delivery; Drive architectural decisions for managed GPU training at scale; Build observability and reliability practices for long-running training jobs; Partner with recruiting to hire top-tier engineering talent

Seniority

Senior, hands-on IC with management responsibilities

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.