DevOps Engineer for Large Scale Compute (IT-CD-CC-2026-223-GRAE)
Core
Manage the lifecycle of 400k+ cores for High Throughput Compute systems, including CI provisioning, automated/AI-enhanced maintenance, and security monitoring.
Role type
Junior DevOps Engineer (Infrastructure & Compute)
Builds
Batch systems (HTCondor, SLURM), HPC systems, and supporting services for the WLCG.
Domain
High Energy Physics / Large Scale Computing Infrastructure
Deliverable
infrastructure
Required skills
Linux administration, Python or Go scripting, Git, Configuration Management (Puppet/Ansible), Infrastructure as Code (Terraform/OpenTofu)
Preferred skills
Container technologies, CI/CD systems, Batch systems (HTCondor, SLURM, kueue)
Responsibilities
Manage fleet lifecycle from provisioning to maintenance, implement LLM-enhanced security monitoring, deploy configuration changes at scale, maintain configuration management infrastructure, develop IaC modules, improve code quality via testing and review.
Seniority
Junior, 0-2 years experience
