(Senior) Cloud Site Reliability Engineer (Scalability) (m/f/x)
Core
Ensure teams can scale up services while being performant, reliable, and cost-optimized through self-service tooling and best-practices for observability, load-testing, and FinOps.
Role type
Senior Cloud Site Reliability Engineer (Scalability)
Builds
Self-service tooling for monitoring, auto-scaling, load testing, and chaos engineering for microservices on AWS.
Domain
Fintech / Cloud Infrastructure / AWS
Deliverable
production ML models | product features | infrastructure
Required skills
AWS, Infrastructure as Code (Terraform), Monitoring, Container Orchestration, Microservices, Python, Distributed Systems Design, Serverless Technologies, CI/CD Automation
Preferred skills
Java/Kotlin, JavaScript/Node.js, GitHub Actions, Jenkins, Chaos Engineering
Technologies
AWS (ECS, Fargate, Lambda), Datadog, Terraform, Python, Java, Kotlin, JavaScript, Node.js, GitHub Actions, Jenkins
Responsibilities
Design and rollout Monitoring best practices in Datadog including SLI, SLO and SLAs; Research and develop service and storage improvements using serverless technologies; Develop and maintain internal tooling around Monitoring, Developer Portal and Load Testing; Mentor and enable software development teams to foster DevOps culture; Design and implement best practices around auto scaling of infrastructure; Run chaos engineering experiments to improve resilience of services.
Seniority
Senior, hands-on IC