Staff ML Infrastructure Engineer (Compute)
Core
Building and scaling robust Compute platforms for Simulation, data labeling, and data generation workflows to support autonomous vehicle validation.
Role type
Staff ML Infrastructure Engineer (Compute)
Builds
Cloud-agnostic, reliable, and cost-efficient infrastructure for ML simulation and hardware-in-loop validation.
Domain
Autonomous Vehicles / High Performance Computing
Deliverable
infrastructure
Required skills
distributed systems design, container orchestration, Go programming, cloud platform management, technical leadership, system architecture, observability, capacity planning, GPU optimization
Preferred skills
Google Compute Engine, hardware-in-the-loop validation, HPC, telemetry integration, hardware acceleration
Technologies
Docker, Kubernetes, Go, GCP, Azure, AWS, Google Compute Engine, GPUs
Responsibilities
Collaborate with Simulation engineers and ML researchers to translate workflows into platform requirements; Own the technical roadmap and lead decisions on Compute architecture, caching, capacity provisioning, and auto-scaling; Drive development of monitoring, observability, and metrics for reliability and resource optimization; Proactively research and integrate frameworks, hardware accelerators, and distributed computing techniques; Lead large-scale technical initiatives across ML infrastructure; Establish best practices to raise the engineering bar.
Seniority
Staff, hands-on IC with strategic influence