Staff Site Reliability Engineer
Core
Lead AI enablement strategy and tooling for the engineering organization while maintaining SRE fundamentals for production infrastructure.
Role type
Staff Site Reliability Engineer (AI Enablement & Platform)
Builds
AI-assisted development tools, frameworks, and guardrails for the engineering org
Domain
Cybersecurity / Insurance / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
SRE fundamentals, AWS, Terraform, Go or Python, AI/LLM-powered tooling, prompt engineering, container orchestration (ECS/Kubernetes), CI/CD, observability
Preferred skills
Distributed systems troubleshooting, event streaming (Kafka/Kinesis), Internal Developer Platforms, systems security, agentic AI workflows, incident response
Technologies
AWS, Terraform, Go, Python, Cursor, GitHub Copilot, ECS, Kubernetes, GitHub Actions, Datadog, Kafka, Kinesis
Responsibilities
Define and drive strategy for embedding AI-native tools into the SDLC; design and develop custom AI tooling and frameworks; partner with teams to automate workflows using agentic tools; establish metrics to measure AI tooling impact; participate in ad-hoc SRE on-call rotation for infrastructure support; mentor engineers and shape best practices
Seniority
Staff, high-influence IC with strategic ownership