AI/MLOps SRE Lead Engineer
Core
Lead the design and operation of resilient AI/ML platforms, LLMs, and cloud-native technologies to support enterprise-scale automation and intelligent systems.
Role type
Senior IC AI/MLOps SRE Lead Engineer
Builds
Enterprise ML platform infrastructure, AI-powered observability solutions, LLM/Agent platforms, and ChatOps integrations
Domain
Healthcare technology, Machine Learning Operations, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, Multi-cloud operations (AWS/GCP/Azure), ML platform technologies (Databricks/SageMaker/Dataiku/Vertex AI), Infrastructure as Code, CI/CD automation, Python/Go/Bash scripting, Observability stack design, Anomaly detection, Predictive analytics, Kubernetes orchestration
Preferred skills
MLOps frameworks (Kubeflow/MLflow), AI Agents/LLM integration, Cloud cost optimization, Policy-as-code, FinOps practices
Technologies
Dataiku, Amazon SageMaker AI, Databricks, Google Vertex AI, Terraform, Pulumi, AWS CDK, Prometheus, Grafana, Datadog, OpenTelemetry, Kubernetes (EKS/GKE/AKS), Kubeflow, Feast, MLflow, LangSmith, RAGAS, Evidently AI
Responsibilities
Establish SLOs/SLIs and reliability methodologies, Design and operate enterprise ML platform infrastructure, Develop AI-powered observability capabilities, Lead implementation of LLM/Agent platforms, Design Infrastructure as Code and self-healing systems, Architect enterprise ChatOps solutions, Mentor engineers and influence platform strategy
Seniority
Senior, hands-on IC with leadership responsibilities