Staff Software Engineer, HPC
Core
Design, scale, and operate custom High-Performance Computing infrastructure to support autonomous vehicle development workflows including AI model training and simulation.
Role type
Staff Software Engineer (HPC Infrastructure)
Builds
Distributed compute infrastructure, job scheduling systems, and developer tools for large-scale workloads
Domain
Autonomous vehicles, High-Performance Computing, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Distributed systems design, Ray.io, Kubernetes, Cloud infrastructure (AWS), Python, System reliability engineering, Capacity planning, Cross-functional leadership
Preferred skills
Machine learning workloads, SLURM, Algorithmic optimization, Operations research, Developer platform engineering
Technologies
Ray.io, Kubernetes, SLURM, AWS, Python
Responsibilities
Design core services and abstractions for distributed compute infrastructure, Build multiyear software engineering roadmap for the HPC platform, Lead cross-team initiatives for org-wide improvements, Create production-grade APIs, SDKs, and tools, Design and improve job scheduling algorithms and auto-scaling policies, Design multi-region orchestration strategies, Resolve systemic reliability and performance issues, Evaluate new technologies for computational capabilities, Develop capacity planning tools and forecasting models, Mentor junior engineers
Seniority
Staff, hands-on IC with strategic scope