CareerPlanGet AI match score →

Staff AI Inference and Acceleration Engineer

HQ💼 Full-time💰 $180,000–$180,000🗓 2026-06-26 → 2026-07-31

Core

Designing and optimizing on-board AI inference architecture for autonomous humanoid robots to meet strict latency, power, and reliability constraints.

Role type

Staff AI Inference and Acceleration Engineer

Builds

On-board inference systems for autonomous humanoid robots

Domain

Robotics, Edge AI, Hardware Acceleration

Deliverable

production ML models

Required skills

AI/ML inference optimization, heterogeneous compute architecture, model quantization and pruning, low-level toolchains (TVM, MLIR, TensorRT), C++ and Python, profiling and benchmarking, memory hierarchy management

Preferred skills

Real-time operating constraints, model-hardware co-design

Technologies

ONNX, TFLite, TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm, NPU, GPU, DSP, CPU

Responsibilities

Map models to accelerators based on latency, power, and memory budgets; Partition inference workloads across heterogeneous compute resources; Define system-level compute budgets; Evaluate next-generation acceleration hardware; Optimize inference toolchains end-to-end; Apply quantization, pruning, and operator fusion; Profile inference pipelines to eliminate bottlenecks; Optimize kernel scheduling and memory layout; Partner with AI/ML teams on model architecture constraints; Work on runtime integration and power management; Engage with silicon vendors to influence hardware roadmaps

Seniority

Staff, hands-on IC with strategic influence

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗