Site Reliability Engineer
Core
Build automation solutions to deploy and maintain applications for 200+ NBC/Telemundo stations, cable networks, and live events (sports, news) serving millions of viewers.
Role type
Site Reliability Engineer (Video Streaming)
Builds
Automation for deploying, managing, and monitoring on-prem and cloud-based broadcast systems and 3rd party vendor integrations.
Domain
Media & Entertainment / Live Broadcast Distribution
Deliverable
production ML models | product features | infrastructure
Required skills
Linux/Unix administration, core networking fundamentals, scripting (Python, Bash, JavaScript, JSON, XML, YML), AWS cloud services, Infrastructure as Code (CloudFormation, Terraform, Ansible, Chef), Git, troubleshooting complex systems, designing automation for issue detection and recovery.
Preferred skills
Containerized systems (Kubernetes, Docker, EKS), Broadcast/Media distribution workflows, Microsoft Graph API, Live Broadcast standards (HEVC, AVC, ATSC, HLS, CMAF, Zixi, SRT, RIST, 2022-7, 2110, SCTE35, SCTE104, SCTE224, SSAI, StatMUX).
Responsibilities
Deploy and maintain systems across facilities and in the Cloud; lead large troubleshooting activities; report bugs to vendors and test new releases; collaborate with operations and engineers to develop solutions; analyze technology and develop improvement processes; develop automation for alerts and recovery; produce documentation for workflows and SOPs; participate in L2 on-call rotation.
Seniority
Mid-level, hands-on IC