AI/MLOps SRE Lead Engineer
Core
Lead the design and operation of resilient AI/ML platforms, ensuring reliability, scalability, and observability for machine learning workloads, LLMs, and AI Agents.
Role type
Senior IC MLOps SRE Lead Engineer
Builds
Enterprise ML platform infrastructure, AI-driven observability solutions, and ChatOps integrations for Data Science and AI Engineering teams.
Domain
Biopharmaceuticals / AI & Machine Learning Platform Engineering
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, Multi-cloud operations (AWS/GCP/Azure), ML platform technologies (Databricks, SageMaker, Dataiku, Vertex AI), Infrastructure as Code (Terraform, Pulumi), CI/CD automation, Python/Go/Bash scripting, Observability (Prometheus, Grafana, OpenTelemetry), Anomaly detection, Kubernetes orchestration.
Preferred skills
MLOps frameworks (Kubeflow, MLflow, Feast), AI Agents/LLM integration, Cloud cost optimization, Policy-as-code.
Responsibilities
Drive service reliability and establish SLOs/SLIs across multi-cloud environments; Design and scale ML platform infrastructure; Develop AI-driven observability and automated remediation; Lead implementation of LLM, RAG, and AI Agent platforms; Architect enterprise ChatOps solutions; Mentor engineers and influence platform strategy.
Seniority
Senior, hands-on IC with leadership responsibilities