Principal Engineer, Compute Fleet Management
Core
Architect and operate the foundational compute infrastructure for Databricks' cloud platform, managing billions of resources across AWS, Azure, and GCP to ensure high availability and efficiency.
Role type
Principal Engineer, Compute Fleet Management
Builds
The core control plane and compute fleet infrastructure enabling Databricks' data and AI platform
Domain
Cloud Infrastructure / Distributed Systems / AI/ML Workloads
Deliverable
infrastructure
Required skills
Large-scale distributed systems design, Cloud platform expertise (AWS/Azure/GCP), Fleet provisioning and pooling, High-availability architecture, Cross-team technical leadership, Strategic project execution
Preferred skills
GPU fleet management for AI/ML, Multi-cloud operational experience
Technologies
AWS, Azure, GCP
Responsibilities
Provision and pool billions of cloud resources for peak performance, Build architecture for horizontal scaling and resilience against cloud failures, Lead development of low-dependency systems for the compute platform, Drive compute utilization efficiency metrics, Enforce security and performance isolation across customer workloads, Manage complex cross-organizational dependencies for strategic initiatives
Seniority
Principal, hands-on IC with strategic leadership