Site Reliability Engineer
Core
Building, managing, maintaining, and securing mission-critical internet infrastructure services including DNS, focusing on Kubernetes environments and physical host operations.
Role type
Senior Site Reliability Engineer (Infrastructure & Kubernetes)
Builds
Mission-critical internet services (DNS) and shared Kubernetes environments
Domain
Internet Infrastructure / Cloud & On-Prem Hybrid
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Linux administration, Kubernetes orchestration, Ansible automation, Infrastructure-as-Code, Python scripting, RHEL 8/9, SELinux, IBM POWER hardware, AIX, OpenStack, Docker, Jenkins, Unix-like systems, network troubleshooting
Preferred skills
ServiceNow, Kanban/Scrum methodologies
Technologies
Kubernetes, Ansible, RHEL, AIX, IBM POWER, OpenStack, Docker, Jenkins, ServiceNow, Python
Responsibilities
Participate in technical designs for migrating to shared Kubernetes environments; Install, configure, and maintain physical hosts for production and non-production applications; Develop deployment automation using Ansible; Orchestrate Kubernetes workloads; Coordinate with technical staff to deploy systems and software; Perform operational support functions including problem isolation and resolution; Document processes, procedures, configurations, and deployment plans; Provide critical technical leadership in Kubernetes environments (on-prem and AWS); Participate in 24x7 on-call rotation; Work with software developers to plan and deploy services.
Seniority
Senior, hands-on IC
