Site Reliability Engineer
Core
Own the stability, availability, and operational health of an agentic AI platform across customer environments.
Role type
Site Reliability Engineer (SRE)
Builds
Cloud infrastructure, deployment pipelines, and observability systems for enterprise AI platforms
Domain
Enterprise Agentic AI / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud platform expertise (GCP, Azure, AWS), Infrastructure-as-Code (Terraform, Ansible), CI/CD pipeline management, Observability tooling (Prometheus, Grafana, Datadog, Splunk), Distributed systems troubleshooting
Preferred skills
Containerization and orchestration (Docker, Kubernetes), High-growth startup experience
Technologies
GCP, Azure, AWS, Terraform, Ansible, GitLab CI, Jenkins, Prometheus, Grafana, Datadog, Splunk, Docker, Kubernetes
Responsibilities
Design and provision cloud infrastructure tailored to customer environments; Execute on-call SaaS deployments with minimal downtime; Monitor logs, alerts, and metrics to maintain SLA commitments; Diagnose and resolve production incidents with root cause analysis; Maintain deployment runbooks and troubleshooting guides
Seniority
Mid-level, hands-on IC