CareerPlanSign in
This role has closed. Browse similar live roles below, or see all open jobs →

ASKUSR0145930 Site Reliability Engineer (SRE)

Berkeley, California, United States💼 Full-time💰 $80–$80🗓 2026-09-03 → 2026-09-25

Core

Maintain accessibility, reliability, security, and operational health of large-scale HPC and data systems for U.S. DOE scientific research.

Role type

Senior IC Site Reliability Engineer (HPC/Data Center)

Builds

24/7 operations environment for mission-critical scientific computing

Domain

High-Performance Computing (HPC) / Data Center Operations

Deliverable

production ML models | infrastructure | physical/clinical work

Required skills

Linux systems administration, incident response, programming/scripting (Python, C, C++, Perl, Java), automation, networking, ServiceNow, physical data center operations, monitoring/alerting

Preferred skills

Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, ITSM best practices, HPC environment support, environmental monitoring (power/cooling), Agentic AI/autonomous automation tools

Technologies

Linux, SSH, Python, C, C++, Perl, Java, ServiceNow, Kubernetes, Prometheus, VictoriaMetrics, Alertmanager, HPC systems, Building management systems

Responsibilities

Monitor HPC systems, storage, networks, and facility infrastructure; respond to alerts and perform initial triage; troubleshoot system, application, and network issues; develop automation solutions to prevent recurring issues; maintain monitoring pipelines and integrations; perform physical and logical walkthroughs of the data center; monitor environmental conditions and power/cooling infrastructure; manage diagnostic software during maintenance; document incidents and operational updates.

Seniority

Senior, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.