Site Reliability Engineer
Core
Ensure reliability, resiliency, and innovation in mission-critical information systems and ecosystems for leading businesses.
Role type
Senior Site Reliability Engineer (SRE)
Builds
End-to-end services spanning customer sites and platforms
Domain
Enterprise IT Operations, Cloud Infrastructure, Multi-cloud (AWS, Azure, GCP)
Deliverable
production ML models | infrastructure
Required skills
Incident management, Application monitoring, Service level objectives (SLOs), Enterprise server management (Windows/Linux/UNIX), Hyperscaler Cloud platforms, Scripting (JSON/YAML/Bash/PowerShell)
Preferred skills
Infrastructure as Code (Ansible/Terraform), Python, Kubernetes, Observability tools (Prometheus/Grafana/Loki)
Technologies
AWS, Azure, Google Cloud Platform, OpenShift, RHEL, AIX, Solaris, Kubernetes, Prometheus, Grafana, Loki, Ansible, Terraform
Responsibilities
Analyze business needs and provide strategic advice for system reliability, Design and implement application monitoring strategies, Manage operations load and overflow using appropriate tooling, Collaborate with DevSecOps and Operations teams on SLOs, Implement cutting-edge tools to enhance operations and security, Identify and mitigate operational issues to ensure seamless customer experiences
Seniority
Senior, hands-on IC