Machine Learning DevOps - Cloud and Compute Cluster - R&D Support
Core
Managing and automating cloud infrastructure for ML training, inference, and production environments to support enterprise AI applications.
Role type
Machine Learning DevOps Engineer (Cloud & Compute Cluster)
Builds
Scalable, reliable ML/LLM pipelines and infrastructure for enterprise AI workloads.
Domain
Artificial Intelligence / Machine Learning Operations / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Linux administration, cluster configuration, containerization and orchestration (Slurm, Docker, Kubernetes), CI/CD workflows, cloud infrastructure (AWS, GCP, Azure), infrastructure as code (Terraform, CloudFormation), ML pipeline orchestration (MLflow, Kubeflow, Airflow, Metaflow), monitoring and logging (Grafana, CloudWatch, Prometheus, Loki), shell scripting, GPU/distributed compute management.
Preferred skills
Experience with terabyte-scale datasets, model versioning and traceability, data drift monitoring.
Technologies
AWS, GCP, Azure, SageMaker Hyperpod, Vertex AI, Slurm, Docker, Kubernetes, GitHub Actions, Jenkins, Gitlab CI, Terraform, CloudFormation, MLflow, Kubeflow, Airflow, Metaflow, Grafana, CloudWatch, Prometheus, Loki.
Responsibilities
Optimize infrastructure for ML training and inference; Automate and maintain ML/LLM pipelines; Manage model versioning, reproducibility, and traceability; Implement ML-centric CI/CD practices; Monitor model performance and data drift in production.
Seniority
Mid-level to Senior, hands-on IC