Senior Site Reliability Engineer (DevOps)
Core
Senior SRE responsible for scaling and hardening production infrastructure on managed Kubernetes and AWS, while building agentic coding workflows and automation to reduce toil.
Role type
Senior Site Reliability Engineer (DevOps)
Builds
Production infrastructure, internal automation tools, CI/CD pipelines, and observability systems
Domain
E-commerce/SaaS tax automation, Cloud Infrastructure (AWS, Kubernetes)
Required skills
AWS (RDS/Postgres, ElastiCache/Redis, VPC), Managed Kubernetes, Infrastructure-as-Code (Terraform, CloudFormation), CI/CD, Observability (metrics, tracing, logging), Agentic coding workflows, Incident response, Cost optimization (via careerplan.io/jobs/e74f7e88-4838-4161-b0b4-a0a14051dedc-senior-site-reliability-engineer-devops-at-kintsugi-ai)
Preferred skills
Multi-cloud experience, Security/Compliance (SOC 2, GDPR), Developer enablement tooling
Technologies
Kubernetes, AWS, Terraform, CloudFormation, Postgres, Redis
Responsibilities
Own reliability of high-traffic production systems, Build and debug using agentic coding workflows, Develop internal tools to eliminate manual toil, Operate monitoring and alerting systems, Partner on reliability design, Automate infrastructure and CI/CD, Lead incident response and postmortems, Optimize infrastructure for cost and security
Seniority
Senior, hands-on IC