Senior Site Reliability Engineer
Core
Own reliability across a cloud-native and AI-driven platform, building systems that scale and self-heal using LLM-powered automation.
Role type
Senior Site Reliability Engineer (AI/LLM focus)
Builds
Cloud-native infrastructure, AI agents for DevOps/SRE workflows, and LLM-powered operational tooling
Domain
Cloud infrastructure, AI/LLM, Kubernetes operations
Deliverable
production ML models | infrastructure
Required skills
AWS infrastructure at scale, Kubernetes (production-grade), distributed systems debugging, Python or Go, monitoring and alerting systems
Preferred skills
LLM APIs (OpenAI), agent frameworks (LangChain, AutoGen), AI agents for DevOps, RAG systems, vector databases, AIOps
Technologies
AWS, Kubernetes, OpenAI API, LangChain, AutoGen, Pinecone, Weaviate, Prometheus, Grafana, ELK, OpenSearch, OpenTelemetry, Terraform
Responsibilities
Own uptime and performance of services on AWS + Kubernetes, design self-healing infrastructure, build LLM-powered tooling for alert triage and incident analysis, manage and scale Kubernetes workloads, build observability systems, define SLOs/SLAs, automate infrastructure, lead incident response, introduce chaos engineering
Seniority
Senior, hands-on IC
