Staff Machine Learning Operations Engineer
Core
Building and maintaining the reliability, performance, and cost-efficiency of production machine learning systems and the platform enabling their deployment.
Role type
Staff Machine Learning Operations Engineer
Builds
Robust ML platform including data infrastructure, model serving, and CI/CD pipelines for model deployment
Domain
Healthcare technology / Machine Learning Infrastructure
Deliverable
infrastructure
Required skills
ML production stack (model serving, feature stores, model registries), Kubernetes, cloud infrastructure (AWS), Infrastructure as Code (Terraform), observability, incident response, CI/CD automation, model optimization
Preferred skills
Healthcare or regulated-data production ML experience, designing ML platforms from scratch
Technologies
Python, Kubernetes, AWS, Sagemaker, Terraform, S3, Snowflake, Airflow, Datadog
Responsibilities
Establish SLOs, observability, and on-call responsibilities for production ML systems; Architect ML platform and data infrastructure; Implement automated ML CI/CD pipelines with data quality checks; Drive down cost and latency through architecture and optimization; Design and implement automated data and concept drift monitoring systems
Seniority
Staff, hands-on IC with strategic leadership