Senior Site Reliability Engineer
Core
Senior SRE building LLM-based agents to automate incident response, root cause analysis, and remediation workflows in production environments.
Role type
Senior IC Site Reliability Engineer (AI/Agentic Automation)
Builds
Autonomous LLM agents, agent tooling, and production reliability infrastructure
Domain
Enterprise AI, Financial Crime Prevention, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
LLM agentic workflows, Python, Kubernetes, cloud provider operations, application-level debugging, observability stack integration, guardrail design, agent evaluation
Preferred skills
Terraform, Ansible, Docker, LLMOps, ML engineering
Technologies
Claude Code, OpenAI Codex, Kimi K2/K3, Kubernetes, AWS/GCP/OCI/Azure, Prometheus, Grafana, Datadog, ELK, PagerDuty, Slack
Responsibilities
Perform hands-on on-call rotation and incident response for workflows without agents; debug application code and ship fixes directly to repos; design and build LLM agents to automate repetitive SRE tasks; define agent tool interfaces and safe wrappers; set guardrails for autonomous actions; build test suites for agent evaluation; tune prompts and tool schemas; report on agent impact metrics.
Seniority
Senior, hands-on IC