Manager, Software Engineering (Resilience Engineering)
Core
Lead the Resilience Engineering team to ensure the safety and reliability of production systems through proactive validation techniques like production load testing and chaos engineering.
Role type
Senior Engineering Manager (Resilience Engineering)
Builds
Platforms and tooling for safe production experimentation, including load testing and fault injection systems.
Domain
Fintech / Distributed Systems / Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Leading engineering teams in reliability/infrastructure, production load testing, chaos engineering, distributed system failure modes, cloud-native environments (AWS, Kubernetes), observability tooling, strong programming (Python, Kotlin, Java), safety guarantees (isolation, rate limiting, guardrails), cross-functional collaboration.
Preferred skills
Experience with chaos engineering vendors (Gremlin, Harness), building reusable tooling/frameworks, influencing engineering practices across teams.
Responsibilities
Define and drive the vision for resilience engineering, lead and mentor engineers building platforms for safe production experimentation, partner with leadership to embed resilience validation in the SDLC, own the design of platforms for load testing and fault injection, establish monitoring and incident response practices for resilience validation, enable teams to adopt resilience practices via tooling and workflows.
Seniority
Senior, hands-on IC with leadership responsibilities