ASKUSR0145930 Site Reliability Engineer (SRE)
Core
Maintain accessibility, reliability, security, and operational health of large-scale HPC and data systems for U.S. DOE scientific research.
Role type
Senior IC Site Reliability Engineer (HPC/Data Center)
Builds
24/7 operations environment for mission-critical scientific computing
Domain
High-Performance Computing (HPC) / Data Center Operations
Deliverable
production ML models | infrastructure | physical/clinical work
Required skills
Linux systems administration, incident response, programming/scripting (Python, C, C++, Perl, Java), automation, networking, ServiceNow, physical data center operations, monitoring/alerting
Preferred skills
Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, ITSM best practices, HPC environment support, environmental monitoring (power/cooling), Agentic AI/autonomous automation tools
Technologies
Linux, SSH, Python, C, C++, Perl, Java, ServiceNow, Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, HPC systems, Building management systems
Responsibilities
Monitor HPC systems, storage, networks, and facility infrastructure; respond to alerts and perform initial triage; troubleshoot system, application, and network issues; develop automation solutions to prevent recurring issues; maintain monitoring pipelines and integrations; perform physical and logical walkthroughs of the data center; monitor environmental conditions and power/cooling infrastructure; manage diagnostic software during maintenance; document incidents and operational updates.
Seniority
Senior, hands-on IC