Machine Learning Ops Engineer
Core
Design, build, and operate the end-to-end AI/ML platform on AWS and Kubernetes to support data scientists and GenAI teams.
Role type
Senior IC Machine Learning Ops Engineer (AI/ML Platform)
Builds
Scalable AI/ML infrastructure, CI/CD pipelines, and observability stacks for production ML models and Agentic AI workloads.
Domain
Healthcare technology, Cloud Infrastructure, Generative AI
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, AWS, Terraform, Python, CI/CD, Platform Observability, AI/ML Platform Engineering, LLM/Agent Workloads
Preferred skills
LangChain, vLLM, RAG, Snowflake, Kubeflow
Technologies
AWS, Kubernetes, Terraform, Prometheus, Loki, Grafana, Datadog, GitHub Actions, LangChain, vLLM, OpenSearch, Snowflake
Responsibilities
Design and operate AI/ML clusters, networking, IAM, and storage on AWS; Provision infrastructure as code with Terraform; Own CI/CD for data pipelines and AI applications; Build observability stacks for metrics, logs, and traces; Automate environment provisioning and self-serve tooling; Partner with teams to enable model serving and safe rollout of Agentic AI; Run POCs for new technologies.
Seniority
Senior, hands-on IC