Senior DevOps Engineer, Infrastructure & Reliabili
Core
Design, build, and operate scalable cloud infrastructure and Kubernetes platforms to ensure high availability and reliability for engineering teams.
Role type
Senior DevOps Engineer (Infrastructure & Reliability)
Builds
Production-grade Kubernetes environments, CI/CD pipelines, and automated infrastructure-as-code systems.
Domain
Cloud Infrastructure, Kubernetes, DevOps
Deliverable
production ML models | product features | infrastructure
Required skills
Terraform, Kubernetes (EKS/self-managed), AWS, CI/CD pipeline optimization, observability (DataDog), disaster recovery planning, incident response leadership, distributed systems architecture, Kafka, PostgreSQL, Redis, service mesh technologies, policy-as-code, multi-region system design, SLOs, error budgets, chaos testing.
Preferred skills
Application development experience, high-throughput Kafka cluster management, internal developer platform experience, zero-trust networking implementation.
Responsibilities
Implement Infrastructure-as-Code patterns for standardized cloud provisioning; Own and evolve the Kubernetes platform; Optimize CI/CD pipelines and improve deployment confidence; Design and enforce secure networking, IAM, and secrets management; Improve observability using metrics, logs, and tracing; Optimize cloud costs through rightsizing and autoscaling; Implement disaster recovery and multi-region resilience; Modernize manual infrastructure into automated systems; Lead incident response and postmortem processes.
Seniority
Senior, hands-on IC