Site Reliability Engineer
Core
Building distributed systems with ultra-low latencies and high throughput for a fraud detection platform serving financial institutions.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Cloud infrastructure, automation tooling, and platforms supporting the fraud detection mission.
Domain
Financial technology / Cloud infrastructure
Deliverable
production ML models | infrastructure
Required skills
Go, Python, distributed systems, asynchronous & multithreaded designs, scalable cloud services, production operations, oncall management, capacity allocation, incident response, root cause analysis, infrastructure as code (IaC), monitoring & alerting, cost optimization.
Preferred skills
Grafana, Prometheus, Kubernetes, Hashicorp, AWS, GCP.
Responsibilities
Provide recommendations on capacity allocation considering cost, resilience, and performance; collaborate with product teams to drive system performance and reliability improvements; develop automation for cloud infrastructure and incident response; create playbooks for actionable alerts; participate in incident response and root cause investigation; maintain and develop infrastructure as code (IaC) for end-to-end lifecycle operations; prevent and investigate production issues.
Seniority
Senior, hands-on IC