Research Engineer - AI/RL Infrastructure
Core
Design, build, and operate large-scale ML infrastructure and training systems to support end-to-end autonomous driving and robotic generalist research.
Role type
Senior/Staff Research Engineer (AI/RL Infrastructure)
Builds
Training and evaluation infrastructure, benchmarking systems, data curation pipelines, and distributed training environments for physical AI.
Domain
Autonomous driving, robotics, physical AI, large-scale distributed systems.
Deliverable
production ML models | infrastructure
Required skills
Large-scale distributed training, performance engineering, compute acceleration, systems-level debugging, open-source ML ecosystem judgment, PyTorch, CUDA, Ray, Flyte, Kubernetes
Preferred skills
Self-driving application experience, Tech Lead or Manager capacity
Technologies
PyTorch, CUDA, Ray, Flyte, Kubernetes
Responsibilities
Design and build training/evaluation infrastructure orchestrating massive GPU clusters; Build robust benchmarking and regression tracking systems; Develop large-scale data sampling and dataset generation pipelines; Enable high-throughput distributed training across heterogeneous cloud environments; Collaborate with research and autonomy teams to translate research into production systems
Seniority
Senior/Staff, hands-on IC with potential for Tech Lead/Manager