Staff MLOps Engineer – ML Platform
Core
Lead the build-out of a cloud-native ML developer platform and production pipelines on AWS to enable teams to move from notebook to secure, reliable, and cost-efficient production services.
Role type
Staff MLOps Engineer (ML Platform)
Builds
Integrated ML/AI development platform with programmatic data analysis and algorithm development capability on AWS
Domain
Cloud Infrastructure / Machine Learning Engineering
Deliverable
production ML models
Required skills
Python, Docker, Terraform, AWS SageMaker, CI/CD for ML, experiment tracking, model registry, data versioning, monitoring & observability, security & compliance
Preferred skills
Distributed training at scale, data engineering at scale, LLMOps/RAG, prior startup experience
Technologies
AWS (SageMaker, S3, IAM, CloudWatch, ECR, ECS/EKS/Lambda), Terraform, Step Functions, Airflow, S3, Glue, EMR/Spark, Athena/Redshift, Great Expectations/Deequ, SageMaker Experiments, MLflow, CodeBuild, CodePipeline, GitHub Actions, FastAPI, Lambda, ECS, EKS, SageMaker Model Monitor, CloudWatch, Prometheus, Grafana, OpenSearch, Bedrock
Responsibilities
Design, build, and operate ML/AI development platform on AWS; Establish golden-path project templates and internal Python libraries; Implement Infrastructure-as-Code and workflow orchestration; Build automated data pipelines with data quality and lineage; Stand up experiment tracking and model registry; Implement CI/CD for ML; Ship real-time and batch endpoints; Build monitoring and observability for production models; Enforce security and governance; Partner with backend engineers to productionize notebooks; Help integrate GenAI/Bedrock services.
Seniority
Staff, hands-on IC with strategic scope