Reliability Engineer
Core
Lead infrastructure resilience and stability for cloud and SaaS environments, implementing AI-driven automation for predictive failure detection and incident resolution.
Role type
Senior Site Reliability Engineer (AI/ML focus)
Builds
Intelligent incident routing, automated service restoration processes, and AI-powered operational assistants.
Domain
Insurance technology, Cloud Infrastructure, AI/ML Operations
Deliverable
production ML models | infrastructure
Required skills
Infrastructure Engineering, Site Reliability Engineering (SRE), DevOps, Observability (Splunk, Dynatrace, CloudWatch), Infrastructure as Code (Terraform, CloudFormation), CI/CD pipeline optimization, Cloud platforms (AWS), Kubernetes, Python, Java, Relational databases (Oracle, SQL Server), Agile methodologies
Preferred skills
AI/ML frameworks for observability, predictive failure detection, LLM-driven troubleshooting, Open-source database technologies
Technologies
Splunk, Dynatrace, CloudWatch, Terraform, CloudFormation, AWS, Kubernetes, Python, Java, Oracle, SQL Server
Responsibilities
Instrument code/application stacks to generate metrics on technology health; Influence technical strategy for architecture rationalization; Develop tooling and alerts to identify and address reliability risks via automation; Enhance delivery flow to increase speed while sustaining reliability; Implement preventative controls and self-healing capabilities; Triage and restore high-impact incidents to minimize downtime; Research and implement AI-based anomaly detection to predict failures; Develop AI-powered troubleshooting copilots and LLM-driven assistants; Implement AI/ML-based runbooks for automated system recovery.
Seniority
Senior, hands-on IC