Senior ML Ops Engineer (Machine Learning Infrastructure)
Core
Design and develop scalable MLOps infrastructure to enable the training, deployment, and monitoring of ML models for autonomous battery-electric rail vehicles.
Role type
Senior IC MLOps Engineer (Machine Learning Infrastructure)
Builds
Scalable ML infrastructure stack for distributed training, experiment tracking, and production deployment
Domain
Autonomous freight transportation / Robotics / Cloud Infrastructure
Deliverable
infrastructure
Required skills
MLOps architecture, distributed training orchestration, CI/CD for ML, cloud platform design, Python, Git, system design
Preferred skills
Deep learning architectures (CNNs, RNNs, Transformers), distributed training tools (PyTorch DDP, Horovod, Ray), real-time ML systems, autonomous vehicles/robotics background
Technologies
MLflow, Kubeflow, SageMaker, Airflow, Metaflow, AWS, GCP, Azure, PyTorch
Responsibilities
Design and implement automated MLOps pipelines for data management, training, and deployment; Architect and manage scalable ML infrastructure for distributed training and inference; Collaborate with ML engineers on data management and deployment strategies; Build and operate cloud-based systems optimized for ML workloads; Support automation of model evaluation and deployment workflows
Seniority
Senior, hands-on IC