Staff AI Inference and Acceleration Engineer
Core
Designing and optimizing on-board AI inference architecture for autonomous humanoid robots to meet strict latency, power, and reliability constraints.
Role type
Staff AI Inference and Acceleration Engineer
Builds
On-board inference systems for autonomous humanoid robots
Domain
Robotics, Edge AI, Hardware Acceleration
Deliverable
production ML models
Required skills
AI/ML inference optimization, heterogeneous compute architecture, model quantization and pruning, low-level toolchains (TVM, MLIR, TensorRT), C++ and Python, profiling and benchmarking, memory hierarchy management
Preferred skills
Real-time operating constraints, model-hardware co-design
Technologies
ONNX, TFLite, TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm, NPU, GPU, DSP, CPU
Responsibilities
Map models to accelerators based on latency, power, and memory budgets; Partition inference workloads across heterogeneous compute resources; Define system-level compute budgets; Evaluate next-generation acceleration hardware; Optimize inference toolchains end-to-end; Apply quantization, pruning, and operator fusion; Profile inference pipelines to eliminate bottlenecks; Optimize kernel scheduling and memory layout; Partner with AI/ML teams on model architecture constraints; Work on runtime integration and power management; Engage with silicon vendors to influence hardware roadmaps
Seniority
Staff, hands-on IC with strategic influence