CareerPlanGet AI match score →

Sr. Engineering Manager, AI Runtime

San Francisco, California💼 Full-time💰 $228,600–$228,600🗓 2026-07-06 → 2026-07-31

Core

Lead the team owning the product experience and foundational infrastructure for Databricks' AI Runtime (AIR), enabling enterprises to train and fine-tune deep learning and LLM models with on-demand GPUs.

Role type

Senior Engineering Manager, AI Runtime Infrastructure

Builds

Managed GPU training infrastructure for deep learning and LLM model development

Domain

Cloud AI Infrastructure / Distributed Deep Learning

Deliverable

production ML models

Required skills

Engineering management, distributed training frameworks (PyTorch, DeepSpeed, Megatron-LM), GPU performance optimization, training resilience patterns, cross-functional leadership, product roadmap definition

Preferred skills

Experience with FSDP, tensor/pipeline parallelism, NCCL, interconnect topologies, memory optimization

Technologies

PyTorch, DeepSpeed, Composer, Megatron-LM, NCCL

Responsibilities

Lead and mentor a high-performing engineering team for Custom Training product and infrastructure; Define and own the product and technical roadmap for AIR; Collaborate with product, research, and platform teams to drive end-to-end delivery; Drive architectural decisions for managed GPU training at scale; Build observability and reliability practices for long-running training jobs; Partner with recruiting to hire top-tier engineering talent

Seniority

Senior, hands-on IC with management responsibilities

Rewrite
## Responsibilities - Lead, mentor, and grow a high-performing engineering team responsible for the Custom Training product and its foundational infrastructure, including distributed training orchestration, cluster lifecycle, fault tolerance, and training efficiency. - Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments. - Collaborate closely with product, research, platform, infrastructure teams, and customers to drive end-to-end delivery, from ideation and prioritization to launch and operation. - Drive architectural decisions and product design for managed GPU training at scale. - Advocate for customer needs through direct engagement, ensuring engineering decisions translate to clear product impact. - Build observability and reliability practices for long-running, multi-node training jobs, including checkpoint strategies, failure recovery, and operational runbooks. - Partner with recruiting to attract, hire, and develop top-tier engineering talent. ## Requirements - 8+ years of software engineering experience, with 3+ years in engineering management. - Track record building and operating managed GPU training infrastructure at scale (100s/1000s GPUs). - Deep familiarity with distributed training frameworks (PyTorch, DeepSpeed, Composer, Megatron-LM) and parallelism strategies (FSDP, tensor/pipeline parallelism). - Experience with training resilience patterns: checkpointing, elastic training, and automated failure recovery for long-running jobs. - Understanding of GPU performance fundamentals including NCCL, interconnect topologies, and memory optimization. - Experience building platform products with clear SLAs where you've owned the customer experience, not just the backend. - Strong cross-functional leadership across platform, product, and research teams, with the ability to lead through ambiguity and deliver complex projects. - Excellent collaboration and communication skills across engineering, product, and research organizations. - BS/MS in Computer Science, Electrical Engineering, or related technical field. ## Nice to Have - None specified. ## Benefits - Competitive salary and total compensation package including annual performance bonus, equity, and benefits. - Opportunities to work with cutting-edge AI and data infrastructure technologies. - Collaborative and innovative work environment with a focus on customer impact and technical excellence. - Global presence with offices around the world. - Access to professional development and career growth opportunities.
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗