GenAI/ML Ops SRE | Python & Multi-Cloud Cloud-Native
Core
Design, deploy, and operate scalable AI-enabled systems and agentic workflows, blending AI/ML deployment with traditional SRE disciplines to ensure reliability, security, and cost-efficiency.
Role type
Senior IC GenAI/ML Ops SRE
Builds
Scalable AI-enabled platforms and agentic workflows
Domain
Artificial Intelligence / Machine Learning Operations / Cloud Infrastructure
Deliverable
production ML models
Required skills
Python, AWS/Azure/GCP, Docker, Kubernetes, CI/CD pipelines, Infrastructure as Code, Observability (Prometheus/Grafana), Incident response, Capacity planning
Preferred skills
MLOps tooling, Multi-cloud strategies, Serverless architectures, Data governance, Regulatory compliance (GDPR/CCPA)
Technologies
Python, AWS, Azure, GCP, Docker, Kubernetes, Terraform, Jenkins, GitHub Actions, GitLab CI, Prometheus, Grafana, CloudWatch, Stackdriver, Git, Jira, Confluence
Responsibilities
Design and operate AI/ML deployment pipelines; Drive automation, observability, and incident response; Collaborate with data scientists and engineers to translate requirements; Define and implement SRE practices and change management; Optimize costs while maintaining SLOs; Lead technical risk assessments and disaster recovery planning; Mentor junior engineers on reliability and security best practices
Seniority
Senior, hands-on IC with mentorship responsibilities