具身智能推理优化工程师
Core
Optimizing inference performance and deployment for embodied intelligence models (multimodal perception, task planning, action strategies) across cloud and edge environments.
Role type
Senior IC embodied intelligence inference optimization engineer
Builds
Low-latency, high-throughput inference services for multi-robot systems
Domain
Robotics + AI Inference Optimization
Deliverable
production ML models
Required skills
C++, Python, PyTorch, model quantization, pruning, distillation, operator fusion, CUDA programming, GPU/NPU architecture knowledge, distributed systems design, monitoring and observability
Preferred skills
Go, TVM, SGLang, vLLM, ONNX Runtime, TensorRT, AI chip toolchains, real-time system engineering
Technologies
PyTorch, TensorRT, vLLM, SGLang, ONNX Runtime, TVM, CUDA, C++, Python, Go
Responsibilities
Profile inference pipelines to identify bottlenecks and validate optimization solutions; Build stable inference service architectures supporting concurrent robot access; Implement monitoring, anomaly detection, and canary release mechanisms; Collaborate with algorithm, hardware, and business teams to bridge lab models to real-world scenarios.
Seniority
Senior, hands-on IC