Machine Learning DevOps - Cloud and Compute Cluster - R&D Support
Core
Manage and automate cloud infrastructure and compute clusters to support ML training, inference, and production pipelines for AI models.
Role type
Senior Machine Learning DevOps Engineer (Cloud & Compute)
Builds
Scalable, automated ML/LLM pipelines and production environments for enterprise AI applications.
Domain
Artificial Intelligence / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
Linux administration, cluster configuration, containerization and orchestration (Slurm, Docker, Kubernetes), CI/CD workflows, cloud infrastructure (AWS, GCP, Azure), infrastructure as code (Terraform), ML pipeline orchestration (MLflow, Kubeflow, Airflow), Python programming, systems and networks administration
Preferred skills
Experience with terabyte-scale datasets, GPU optimization, distributed compute management, monitoring and data drift detection
Technologies
Kubernetes, Docker, Slurm, Terraform, AWS SageMaker, GCP Vertex AI, GitHub Actions, Jenkins, Grafana, Prometheus, MLflow, Kubeflow, Airflow, PyTorch, TensorFlow
Responsibilities
Optimize infrastructure for ML training and inference; Automate and maintain ML/LLM pipelines; Manage model versioning, reproducibility, and traceability; Implement ML-centric CI/CD practices; Monitor model performance and data drift in production
Seniority
Senior, hands-on IC