AI DevOps Engineer (MLOps & Cloud)
Core
Design, implement, and maintain scalable, secure cloud infrastructure for AI/ML solutions, supporting the end-to-end MLOps lifecycle from training to production deployment.
Role type
Senior IC MLOps and Cloud Infrastructure Engineer
Builds
Production-ready AI/ML systems, automated deployment pipelines, and secure cloud estates for global enterprises
Domain
Enterprise technology, AI/ML, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
MLOps lifecycle management, Infrastructure as Code (Terraform, CloudFormation), CI/CD pipeline development, container orchestration (Kubernetes, Docker), cloud platform expertise (AWS, Azure, GCP), scripting and automation (Python, Bash), observability and monitoring (Prometheus, Grafana, ELK, Datadog), DevSecOps practices, secrets management (Vault, AWS Secrets Manager)
Preferred skills
GPU workload optimization, AI-assisted tooling usage, agile environment adaptability
Technologies
Terraform, CloudFormation, GitHub Actions, GitLab CI, Jenkins, Azure DevOps, MLflow, Kubeflow, Weights & Biases, SageMaker Pipelines, Docker, Kubernetes, Helm, Prometheus, Grafana, ELK, Datadog, Vault, AWS Secrets Manager, Python, Bash
Responsibilities
Design and maintain scalable cloud infrastructure for AI/ML solutions; Build and manage Infrastructure as Code; Develop and maintain CI/CD pipelines for AI applications; Support end-to-end MLOps lifecycle; Automate deployments using best practices; Configure monitoring, alerting, and observability; Optimize performance, scalability, and cost; Troubleshoot production issues; Support developers and data scientists on DevOps tooling; Implement DevSecOps standards; Maintain technical documentation
Seniority
Senior, hands-on IC