Sr. Software Engineer – Cloud Infrastructure and Devops
Core
Design, build, and operate infrastructure, tooling, and processes to ensure Venmo's critical systems can withstand and rapidly recover from any failure scenario, including disaster recovery automation, chaos engineering, and incident response.
Role type
Senior Cloud Infrastructure and DevOps Engineer (Business Continuity)
Builds
Disaster recovery automation, chaos engineering experiments, backup and restore tooling, and resilience runbooks for Venmo's AWS cloud environment.
Domain
Fintech / Payments / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
AWS (IaaS/PaaS), Kubernetes (EKS), Docker, Terraform, Python, Go, Bash, Chaos Engineering, Disaster Recovery, SLOs/RTO/RPO
Preferred skills
AWS Fault Injection Simulator, Gremlin, Chaos Monkey, Litmus, Game Day exercises, Incident management
Technologies
AWS, EKS, Docker, GitHub Enterprise, Terraform, GitHub Actions, DataDog, Bash, Python, Go
Responsibilities
Design and implement disaster recovery automation and backup/restore processes; Conduct chaos engineering experiments and incident game days; Troubleshoot incidents, identify root causes, and implement preventive measures; Mentor junior engineers and drive best practices; Define and track resilience metrics, SLOs, and recovery objectives; Develop tools and automation for infrastructure as code.
Seniority
Senior, hands-on IC with mentorship responsibilities