Engineering Manager, Fleet Engineering
Core
Lead and grow distributed engineering teams responsible for the full lifecycle of production GPU fleet infrastructure, including deployment, reliability, orchestration, and foundation tooling.
Role type
Engineering Manager, Infrastructure (GPU Fleet)
Builds
Production-ready GPU clusters, fleet orchestration systems, and host enablement tooling for AI cloud infrastructure.
Domain
AI/ML Cloud Infrastructure, Datacenter Operations, Hardware-Software Integration
Deliverable
production ML models | infrastructure
Required skills
Team leadership and growth, cross-functional project delivery, Linux systems debugging, technical design for medium-to-large efforts, stakeholder management, incident management, hiring and performance management
Preferred skills
Linux systems administration, TCP/IP networking, bare metal provisioning, automation scripting, distributed systems, GPU acceleration, datacenter physical infrastructure, AI-assisted development tools
Technologies
InfiniBand, Linux, PXE, Redfish, IPMI, BMC, DHCP, DNS, NetBox, GPUs, Virtualization
Responsibilities
Lead and grow a distributed team of engineers, deliver projects and deployments on time, identify efficiency gains in tools and processes, provide visibility into project progress and risks, participate in technology qualification, manage staff allocation and priorities, conduct 1:1s and support career development, contribute to incident management and reviews
Seniority
Manager, hands-on leadership