Associate, Site Reliability Engineer (Platform Support), SRE and Governance, Group Technology
Core
Administer the full Kubernetes and OpenShift platform life cycle, ensuring security, reliability, and high availability while managing infrastructure for tenant applications.
Role type
Associate Site Reliability Engineer (Platform Support)
Builds
Secure, reliable, and highly available Kubernetes/OpenShift infrastructure platforms
Domain
Cloud Infrastructure / Container Orchestration / System Administration
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes orchestration, OpenShift administration, Docker containerization, Linux OS management, Windows Server administration, server hardening, patching, vulnerability remediation, incident management, capacity management, monitoring and alerting, configuration management (Ansible, Terraform, Chef), troubleshooting OS/application/network issues, operational documentation
Preferred skills
Production experience with containers (Docker, OpenShift, Kubernetes), experience with HP and DELL hardware (Blades and Rack servers)
Technologies
Kubernetes, OpenShift, Docker, Linux, Windows Server, Jira, Jenkins, Ansible, Terraform, Chef, IIS, SSL, OpenSSH, tectia, IBM CD
Responsibilities
Administer Kubernetes platform life cycle, support OpenShift infrastructure and resolve tenant application issues, automate platform operations, configure monitoring and alerts, manage capacity, perform server hardening and patching, troubleshoot OS and application issues, respond to incidents and production changes, maintain operational documentation and SOPs
Seniority
Associate, hands-on IC