R&D-023 MLOps Engineer
Core
Designing and maintaining large-scale ML pipelines and distributed training clusters for training Vision-Language-Action (VLA) models on massive humanoid robot datasets.
Role type
Senior MLOps Engineer (Robotics & Foundation Models)
Builds
Distributed training clusters, MLOps tools, and evaluation pipelines for VLA models
Domain
Robotics, Embodied AI, Large-scale Machine Learning
Deliverable
production ML models
Required skills
MLOps engineering, distributed systems, cloud services (AWS/GCP/Azure), orchestration tools (Airflow/Dagster/Kedro), PyTorch or JAX
Preferred skills
RL/VLM/VLA training techniques, ETL/ELT pipelines, hyper-parameter optimization, distributed training frameworks, GPU memory management, robotics sensor data processing (RGB/Depth/point clouds), ROS/ROS2, SQL, performance analysis tools
Technologies
PyTorch, JAX, Airflow, Dagster, Kedro, AWS, GCP, Azure, ROS, ROS2
Responsibilities
Design and implement large-scale ML pipelines for robot data; Deploy and maintain distributed training clusters; Collaborate with researchers on ML infrastructure and data pre-processing; Optimize ML infrastructure for cost, performance, and reliability; Develop MLOps tools for performance visualization and analysis
Seniority
Senior, hands-on IC