Chaos Test Lead
Core
Lead the execution of chaos engineering lifecycle to ensure system resilience and high availability by designing, automating, and reporting on failure experiments.
Role type
Senior IC chaos engineering lead
Builds
Resilient distributed systems and automated chaos experiments
Domain
Cloud infrastructure and distributed systems reliability
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Chaos engineering lifecycle management, automated experiment design, distributed system architecture analysis, CI/CD pipeline troubleshooting, Unix/Linux OS internals, public cloud platforms (AWS/GCP/Azure), monitoring and alerting, VPC and load balancer configuration, complex distributed system debugging, chaos engineering tools (Gremlin/Litmus/Chaos Native)
Preferred skills
Leadership in cross-functional environments, stakeholder collaboration
Technologies
Gremlin, Litmus, Chaos Native, AWS, GCP, Azure, Unix/Linux, CI/CD pipelines
Responsibilities
Implement and lead execution of the chaos engineering Lifecycle - Chaos Test Planning, Chaos Test Designing, and Reporting; Ensure recovery and resilience testing is scheduled, staffed, executed, and documented, including remediation and closure of issues; Analyse architecture and recommend weak areas likely to cause failure or outages; Design, develop and execute automated / continuous Chaos Engineering experiments; Troubleshoot failures in CI/CD pipeline; Automate Chaos experiments through chaos engineering tools to run continuously
Seniority
Senior, hands-on IC
