AI DevOps & Reliability Engineer
Core
Own software delivery and operational reliability while leading the adoption of AI tooling (Claude Code, agentic workflows) into runbook generation, alerting, and incident response.
Role type
Senior IC DevOps & Reliability Engineer (AI-embedded)
Builds
CI/CD pipelines, ephemeral environments, GitOps delivery, and AI-augmented operational tooling
Domain
SaaS / Mobile Growth Attribution / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, AWS, Infrastructure as Code (Terraform/CloudFormation), CI/CD architecture, GitOps (Argo CD), streaming infrastructure (Kafka), SQL/NoSQL management, Python/Bash scripting, observability stacks (Prometheus/Grafana/PagerDuty), incident response, SLI/SLO definition
Preferred skills
Progressive delivery (canary/blue-green), cloud cost optimization, Spark/schema management
Technologies
Kubernetes, AWS, Terraform, CloudFormation, Argo CD, Kafka, Prometheus, Grafana, PagerDuty, Python, Bash, Claude Code
Responsibilities
Design and expand deployment automation for on-demand releases; establish release practices and standards; build self-service ephemeral environments; integrate AI tooling into operations; champion Infrastructure as Code; drive GitOps-based delivery; manage high-volume data infrastructure; mentor embedded teams on operational best practices; define and track DORA metrics
Seniority
Senior, hands-on IC with mentorship responsibilities