Site Reliability Engineer
Core
Build internal tools and automation to monitor, diagnose, and support a fleet of client deployments across AWS and Azure, ensuring standardization and observability.
Role type
Site Reliability Engineer (SRE)
Builds
Internal automation tools, monitoring dashboards, and deployment pipelines for client environments
Domain
Financial services / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python, AWS or Azure, Terraform, Linux, Infrastructure as Code, Observability, Incident Management
Preferred skills
Multi-tenant environment operations, Financial services domain knowledge
Technologies
Python, AWS, Azure, Terraform, Linux
Responsibilities
Build internal tools and automation in Python to monitor and support client deployments; Drive standardization across client environments to detect and remediate configuration drift; Improve fleet-wide observability by building monitoring and alerting systems; Convert runbooks into automated checks and self-healing jobs; Extend client provisioning and deployment pipelines; Work with client-facing teams to identify operational toil
Seniority
Mid-level, hands-on IC