SRE Engineer
Core
Define, monitor, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets for AI services while maintaining scalable, fault-tolerant infrastructure for AI model serving, training pipelines, and data workflows.
Role type
Senior Site Reliability Engineer (AI/MLOps)
Builds
Production-grade AI microservices, ML workflows, and data pipelines on Azure
Domain
Industrial software, Artificial Intelligence, Cloud Infrastructure
Deliverable
production ML models
Required skills
SLO/SLI/Error Budget management, Incident response and RCA, Azure cloud services, Kubernetes, Docker, Terraform, CI/CD pipelines, Python, Bash, Go, Observability stacks (Prometheus, Grafana, Datadog, Open Telemetry), MLOps practices
Preferred skills
Supervising and logging tools (ELK Stack), Zero-trust architecture principles, Industry compliance standards (SOC 2, ISO 27001, GDPR)
Technologies
Azure, AKS, Azure DevOps, ARM, Terraform, Ansible, Jenkins, GitHub Actions, Prometheus, Grafana, Datadog, Open Telemetry, Docker, Kubernetes, Python, Bash, Go
Responsibilities
Enforce SLOs/SLIs/Error Budgets for AI services; Lead incident response and root cause analysis; Proactively identify and resolve reliability risks; Maintain scalable infrastructure for data pipelines and ML workflows; Handle cloud infrastructure using IaC tools; Oversee Kubernetes clusters and containerized workloads; Maintain observability stacks; Develop intelligent alerting systems; Automate operational tasks via scripting; Build and maintain CI/CD pipelines; Implement MLOps practices; Ensure alignment with security guidelines; Partner with data engineers and scientists; Foster reliability-first culture; Contribute to on-call rotations
Seniority
Senior, hands-on IC