Inference
Core
Build low-latency inference pipelines for on-device deployment and design distributed inference systems on GPU clusters for robotics applications.
Role type
Senior IC machine-learning infrastructure engineer (inference)
Builds
Low-latency inference pipelines, distributed GPU serving systems, and monitoring/debugging tools for robotics
Domain
Robotics + High-performance ML inference
Deliverable
production ML models
Required skills
Distributed systems, Python, C++/Rust/Go, CUDA, Triton, kernel optimization, quantization, memory management, compute scheduling, graph compilation
Preferred skills
Experience scaling inference workloads in cluster and on-device environments, hardware–software tuning expertise
Technologies
CUDA, Triton, Python, C++, Rust, Go
Responsibilities
Build low-latency inference pipelines for on-device deployment; Design and optimize distributed inference systems on GPU clusters; Implement efficient low-level code and integrate into high-level frameworks; Optimize workloads for throughput and latency; Develop monitoring and debugging tools for reliability and determinism
Seniority
Senior, hands-on IC