Senior DevOps Engineer, Infrastructure & Reliability
Core
Build and maintain scalable, reliable cloud infrastructure and Kubernetes platforms to enable engineering teams to ship software faster and more safely.
Role type
Senior hands-on IC DevOps/SRE engineer
Builds
Production-grade cloud infrastructure, Kubernetes clusters, CI/CD pipelines, and observability systems
Domain
Cloud infrastructure (AWS) and container orchestration
Deliverable
production ML models | product features | infrastructure
Required skills
Terraform, Kubernetes, AWS, CI/CD pipeline design, observability, disaster recovery planning, cloud cost optimization, distributed systems, event-driven architectures, database performance tuning
Preferred skills
Application coding, high-throughput Kafka cluster operation, autoscaling strategies, service mesh technologies, internal developer platforms, zero-trust networking, multi-region distributed systems, reliability frameworks (SLOs, chaos testing)
Technologies
Terraform, Kubernetes, ArgoCD, GitHub Actions, DataDog, AWS (EKS, RDS, MSK, S3, Lambda, IAM, VPC), PostgreSQL, Kafka, Redis, Bash, Python, TypeScript, JavaScript
Responsibilities
Implement scalable Infrastructure-as-Code patterns, own and evolve the Kubernetes platform, optimize CI/CD pipelines, design secure networking and secrets management, improve observability, optimize cloud costs, implement disaster recovery, refactor manual infrastructure into automated systems, introduce new tooling, partner with engineering teams to reduce friction
Seniority
Senior, hands-on IC