ML Platform Engineer (m/f/d)
Core
Build and operate distributed training, deployment, and experimentation infrastructure for research, data, and robotics teams to move models from prototype to production.
Role type
Senior ML Platform Engineer
Builds
Distributed training workflows, containerized ML environments, CI/CD pipelines, and observability systems for ML workloads
Domain
Robotics, Artificial Intelligence, Cloud Infrastructure
Deliverable
infrastructure
Required skills
Distributed training systems, Kubernetes, Docker, PyTorch Distributed, DeepSpeed, CI/CD pipelines, Cloud infrastructure (AWS), Experiment tracking, Model versioning, Observability, Python, System design
Preferred skills
Multimodal systems, Infrastructure as code, ML orchestration, Distributed compute
Responsibilities
Design and scale distributed training workflows for large models; Build and maintain containerized ML environments; Develop and maintain CI/CD pipelines for ML systems; Implement experiment tracking and model versioning workflows; Set up monitoring systems to track model performance and detect drift; Collaborate with research, data, and robotics teams to connect models to production systems