DevOps & AI/ML Infrastructure Engineer
Core
Design, deploy, operate, and secure cloud and ML/AI infrastructure, automating deployments and maintaining CI/CD pipelines for scalable, secure development workflows.
Role type
Senior DevOps & AI/ML Infrastructure Engineer
Builds
Scalable cloud environments, CI/CD pipelines, and production ML platform infrastructure for AI/ML products and agentic use cases.
Domain
Cloud Infrastructure, MLOps, AI/ML Platform Engineering
Deliverable
production ML models | infrastructure
Required skills
Cloud infrastructure management, Infrastructure as Code (Terraform, Terragrunt, CloudFormation), CI/CD pipeline maintenance, Kubernetes orchestration, Linux system administration, Python/Bash scripting, Networking fundamentals, AI tooling integration, MLOps platform operations, Observability and monitoring.
Preferred skills
Google Cloud experience, Helm and service mesh technologies, Serverless and event-driven architectures, Cloud security practices, FinOps, API management platforms, MLOps platforms (Databricks, model serving).
Technologies
AWS (EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway), GitLab CI/CD, Jenkins, Terraform, Terragrunt, CloudFormation, Kubernetes, Amazon EKS, Databricks, Prometheus, Grafana, Coralogix, CloudWatch, Helm, Istio, Linkerd, Traefik, Kong, Apigee, Nessus, Prowler, Trivy.
Responsibilities
Provision and manage cloud resources using IaC; Implement cloud security best practices including IAM and vulnerability management; Maintain and optimize CI/CD pipelines for application and ML model deployments; Operate and scale ML platform infrastructure including Databricks clusters and model serving endpoints; Maintain monitoring, logging, metrics, and alerting solutions; Support incident response and perform Root Cause Analysis; Collaborate with ML Engineering on model deployment and quality thresholds.
Seniority
Senior, hands-on IC