CareerPlanSign in

Inference

Bay Area💼 Full-time🗓 2026-04-30 → 2026-09-25

Core

Build low-latency inference pipelines for on-device deployment and design distributed inference systems on GPU clusters for robotics applications.

Role type

Senior IC machine-learning infrastructure engineer (inference)

Builds

Low-latency inference pipelines, distributed GPU serving systems, and monitoring/debugging tools for robotics

Domain

Robotics + High-performance ML inference

Deliverable

production ML models

Required skills

Distributed systems, Python, C++/Rust/Go, CUDA, Triton, kernel optimization, quantization, memory management, compute scheduling, graph compilation

Preferred skills

Experience scaling inference workloads in cluster and on-device environments, hardware–software tuning expertise

Technologies

CUDA, Triton, Python, C++, Rust, Go

Responsibilities

Build low-latency inference pipelines for on-device deployment; Design and optimize distributed inference systems on GPU clusters; Implement efficient low-level code and integrate into high-level frameworks; Optimize workloads for throughput and latency; Develop monitoring and debugging tools for reliability and determinism

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.