HPC Data Center Developer
Core
Build and own automation and tooling for HPC data center operations, focusing on hardware onboarding, lifecycle management, capacity planning, outage simulation, and monitoring integration.
Role type
Senior IC infrastructure automation engineer (HPC data center)
Builds
Production-ready automation systems, capacity planning tools, outage simulation models, and monitoring integrations for HPC facilities
Domain
High-Performance Computing (HPC) infrastructure and data center operations
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Golang, Linux systems administration, hardware provisioning automation, data center infrastructure knowledge (power/cooling/cabling), IPMI/BMC/Redfish/SNMP integration, observability stack (Grafana/Prometheus/InfluxDB), infrastructure-as-code (Terraform/Ansible), networking (BGP/VLANs), SQL (ClickHouse/MySQL), CI/CD workflows, AI tool usage
Preferred skills
Python, experience with large-scale data center environments, root cause analysis expertise
Technologies
Golang, Python, Linux, Grafana, Prometheus, InfluxDB, SaltStack, Ansible, Terraform, ClickHouse, MySQL, GitHub, IPMI, BMC, Redfish, SNMP, Arista, Cisco
Responsibilities
Design and maintain automation for hardware onboarding and lifecycle management; develop tools for power/cooling capacity planning and outage simulation; build monitoring integrations and alerting strategies; collaborate with operations leads to translate requirements into production systems; maintain reliability and documentation for all developed tools; participate in coordinated maintenance operations
Seniority
Senior, hands-on IC
