ML Ops Enigneer / 1
Core
Design and implement infrastructure to host, orchestrate, and manage up to 1,500 ML scoring processes within a Databricks environment.
Role type
ML Ops Engineer
Builds
Scalable, secure, and well-monitored platforms for data science teams to deploy ML models.
Domain
Cloud infrastructure / Machine Learning Operations
Deliverable
infrastructure
Required skills
Databricks cluster management, Terraform (Infrastructure as Code), CI/CD pipeline integration, job orchestration, monitoring and alerting, model versioning (MLflow), resource optimization, fault tolerance design.
Preferred skills
Experience with Delta Lake, cloud-native tools, collaboration with DevOps teams.
Technologies
Databricks, Terraform, MLflow, Delta Lake, CI/CD tools.
Responsibilities
Set up Databricks clusters, jobs, and workflows for large-scale ML scoring; implement scalable infrastructure for thousands of ML scoring tasks; configure job scheduling and parallel execution strategies; integrate logging, alerting, and dashboards for monitoring scoring throughput and latency; automate resource provisioning and deployments from CI/CD pipelines; define operational SLAs for scoring workloads.
Seniority
Mid-level to Senior, hands-on IC