Staff IT Site Reliability Engineer (Core - Linux)
Core
Design and maintain resilient hybrid cloud infrastructure and self-healing systems using AI-driven automation and predictive modeling to ensure global service availability.
Role type
Staff Site Reliability Engineer (Linux)
Builds
Production-grade hybrid cloud environments, automated incident response playbooks, and AI-integrated monitoring systems.
Domain
Cybersecurity infrastructure, Linux systems, Cloud Orchestration
Deliverable
production ML models | infrastructure
Required skills
Linux kernel tuning, Ansible, Terraform, GCP/GKE/AWS, P1/P2 incident leadership, automated patching, API development, CIS compliance
Preferred skills
AI coding assistants (GitHub Copilot/Claude), Generative AI for documentation, intent-based infrastructure management
Technologies
RHEL, Ubuntu, Ansible, Chef, Terraform, GCP, GKE, AWS, Azure, GitHub Copilot, Claude
Responsibilities
Configure resilient hybrid cloud deployments, manage capacity planning with AI modeling, lead P1/P2 incident response, design self-healing systems, deploy automation frameworks, design proactive monitoring systems
Seniority
Staff, hands-on IC with strategic scope