Lead Site Reliability Engineer
Core
Lead SRE driving reliability, scalability, and performance for critical Disney Experiences platforms powering theme parks, resorts, cruises, and consumer apps.
Role type
Senior IC lead Site Reliability Engineer
Builds
Cloud-native and hybrid infrastructures, CI/CD pipelines, observability systems, and automation for mission-critical guest-facing applications.
Domain
Entertainment technology, large-scale distributed systems, cloud infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud architecture (AWS/Azure/GCP), Kubernetes/Docker/ECS/AKS/GKE, Infrastructure as Code (Terraform/Ansible), CI/CD (Jenkins/GitLab/Azure DevOps), Python/NodeJS/Golang/Bash, UNIX/Linux administration, Observability (ELK/Datadog/Splunk), Networking (TCP/IP/DNS/TLS), Database management (MySQL/MongoDB/DynamoDB/Redis), Incident response, Mentoring
Preferred skills
AI/automation for predictive insights, Java/NodeJS/Python in large-scale environments, Multi-origin hybrid architecture design
Technologies
AWS, Azure, Google Cloud, Docker, Kubernetes, ECS, AKS, GKE, Terraform, CloudFormation, Ansible, Chef, GitHub, GitLab, Jenkins, AWS CodeBuild, Azure DevOps, AppDynamics, New Relic, ELK stack, Datadog, Splunk, MySQL, MongoDB, DynamoDB, Redis, Snowflake, Tableau
Responsibilities
Architect scalable cloud-native and container-based infrastructure; Lead DevOps/SRE practices including CI/CD design and automation; Define reliability strategies, SLIs/SLOs/SLAs, and drive uptime improvements; Lead major incident response and root cause analysis; Develop Infrastructure as Code and automation scripts; Collaborate on capacity planning, security, and migration strategies; Mentor engineers and guide technical direction.
Seniority
Senior, hands-on IC with leadership responsibilities