Senior Infrastructure Engineer, SRE
Core
Lead the reliability and operational evolution of a high-scale cloud platform processing billions of transactions and hundreds of millions of logs daily.
Role type
Senior Staff Site Reliability Engineer (SRE) / Cloud Infrastructure Engineer
Builds
Cloud infrastructure, observability platform, disaster recovery strategies, and reliability tooling
Domain
Fintech / Cloud Infrastructure / Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud infrastructure engineering, SRE practices, SLI/SLO definition, disaster recovery planning, observability platform management, Python/Go/TypeScript, Terraform, AWS
Preferred skills
Reliability modernization leadership, internal tooling development, chaos engineering, observability cost optimization
Technologies
Datadog, Terraform, AWS, Python, Go, TypeScript
Responsibilities
Build and improve system reliability and resiliency; Establish and review SLIs, SLOs, and error budgets; Own and evolve disaster recovery strategy; Partner with product engineering teams on service ownership; Evolve observability platform standards; Strengthen incident practice and postmortem actions; Contribute to day-to-day cloud infrastructure build-outs and on-call rotation
Seniority
Senior, hands-on IC with substantial team impact