Staff Machine Learning Operations Engineer
Core
Build and maintain the reliability, performance, and cost-efficiency of production machine learning systems and the platform enabling their deployment.
Role type
Staff Machine Learning Operations Engineer (Platform Engineering)
Builds
Robust ML platform including data infrastructure, model serving, and CI/CD pipelines for secure model deployment
Domain
Healthcare / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
ML production stack (model serving, feature stores, model registries), Kubernetes, cloud infrastructure (AWS), Infrastructure as Code (Terraform), observability, incident response, CI/CD pipeline design, model optimization
Preferred skills
Healthcare or regulated-data production ML experience, building ML platforms from scratch
Technologies
Python, Kubernetes, AWS, Sagemaker, Terraform, S3, Snowflake, Airflow, Datadog
Responsibilities
Establish SLOs, observability, and on-call responsibilities for ML systems; Architect ML platform and data infrastructure; Implement automated ML CI/CD pipelines with data quality checks; Drive down cost and latency via architecture and optimization; Design and implement automated data and concept drift monitoring systems
Seniority
Staff, hands-on IC with strategic leadership