Site Reliability Engineer
Core
Design, develop, and implement automated CI/CD pipelines and infrastructure-as-code solutions to ensure the reliability, availability, and scalability of mission-critical data management and migration platforms.
Role type
Senior Site Reliability Engineer (Data Management/Migration)
Builds
Automated deployment pipelines, cloud infrastructure, and network configurations for data platforms
Domain
Financial Services / Cloud Infrastructure & Data Management
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering (SRE) fundamentals, CI/CD pipeline design, Infrastructure as Code (IaC), Cloud infrastructure management, Observability and monitoring, Incident response, Capacity planning, SLO/SLI definition and tracking, Python programming, Container orchestration
Preferred skills
Experience with enterprise AI tools for incident triage, Troubleshooting complex networking issues, Toil reduction strategies
Technologies
Ansible, CI/CD, Datadog, Docker, Dynatrace, GitLab, Grafana, Jenkins, Kubernetes, Prometheus, Splunk, Terraform, Linux, Windows
Responsibilities
Guide colleagues in creating solution designs and building alignment around SRE best practices, Design and implement deployment and reliability approaches through automated CI/CD pipelines, Build infrastructure, configuration, and network as code for applications and platforms, Use enterprise-authorized AI capabilities to speed up incident triage and post-incident analysis, Partner with technical specialists to resolve complex issues and address risks proactively using SLIs and SLOs, Improve availability, reliability, and scalability of applications by iterating on solutions