ML Ops Engineer
Core
Own the infrastructure and operational lifecycle of machine learning systems powering a clinical monitoring platform to identify early signs of clinical deterioration.
Role type
Senior ML Ops Engineer
Builds
Production ML pipelines, deployment infrastructure, and monitoring systems for predictive models
Domain
Healthcare technology / Machine Learning Operations
Deliverable
production ML models
Required skills
Python, Apache Airflow, MLflow, AWS (Batch, EC2, S3, IAM, CloudWatch), Docker, Infrastructure-as-Code, Git, SQL, Snowflake, data drift detection, CI/CD for models
Preferred skills
Edge/embedded device deployment, model serving frameworks (TorchServe, TF Serving, Triton), data versioning tools (DVC, LakeFS), distributed compute (Apache Spark, Dask), streaming architectures
Technologies
Apache Airflow, MLflow, AWS Batch, AWS EC2, AWS S3, AWS IAM, AWS CloudWatch, Docker, Snowflake, GitHub Actions, Jenkins, TorchServe, TF Serving, Triton, DVC, LakeFS, Apache Spark, Dask
Responsibilities
Own and extend ML pipeline orchestration; Build automated pipelines for model retraining and validation; Implement pipeline monitoring and failure recovery; Design architectures for rapid experimentation and reproducibility; Deploy and manage ML models on AWS and edge devices; Manage model versioning and rollback workflows; Implement safe model rollout strategies; Build monitoring systems for model performance and data drift; Develop incident response procedures; Manage and optimize AWS compute resources; Design infrastructure-as-code solutions; Drive cost optimization; Support Snowflake integrations; Introduce ML engineering best practices; Build internal tooling for development-to-production cycles; Document operational processes; Participate in architecture discussions; Ensure compliance with healthcare security and privacy requirements (HIPAA, SOC 2)
Seniority
Senior, hands-on IC