Site Reliability Engineer - DevSecOps Engineer
Core
Ensuring reliability, resiliency, and innovation in mission-critical information systems and ecosystems for the world's leading businesses.
Role type
Senior Site Reliability Engineer (DevSecOps)
Builds
End-to-end services spanning customer sites and platforms
Domain
Enterprise IT, Cloud Infrastructure, DevSecOps
Deliverable
production ML models | infrastructure
Required skills
Incident management, Application monitoring, Service level objectives definition, Linux server administration, Networking, Storage, Hyperscaler Cloud platforms, Data format handling, Scripting languages
Preferred skills
Infrastructure as Code (Ansible, Terraform), Python, Container orchestration (Kubernetes, OpenShift), Observability tooling (Prometheus, Grafana, Loki)
Technologies
RHEL, AWS, Azure, Google Cloud Platform, OpenShift, Kubernetes, Prometheus, Grafana, Loki, Ansible, Terraform, Python, Bash, JSON, YAML
Responsibilities
Analyze business needs and provide strategic advice and designs for system reliability, Implement strategies to cap operations load and handle overflow, Collaborate with stakeholders to define service level indicators and objectives, Work alongside development and operations teams to maintain robust systems, Identify and mitigate common operational issues, Deploy changes and maintain systems throughout the software lifecycle
Seniority
Senior, hands-on IC