Software Engineer, HPC Scheduling
Core
Designing and developing high-quality software solutions for a large high-performance compute (HPC) platform to enable complex research at scale.
Role type
Senior IC software engineer (HPC scheduling & Kubernetes)
Builds
Multi-cluster Kubernetes batch job scheduling systems (Armada) and globally distributed infrastructure for research workloads
Domain
High-performance computing, cloud infrastructure, and batch processing
Deliverable
production ML models | infrastructure
Required skills
Golang, Kubernetes (controllers/operators), event-driven programming, Apache Kafka, Apache Pulsar, DAG workflows, PostgreSQL, Linux system administration, networking, CI/CD pipelines
Preferred skills
Armada, SLURM, AWS, Prometheus, Grafana
Technologies
Golang, Kubernetes, Apache Kafka, Apache Pulsar, PostgreSQL, Prometheus, Grafana, Armada, SLURM, AWS
Responsibilities
Designing and developing software solutions using procedural programming languages; Building and maintaining highly scalable, globally distributed systems; Managing and optimizing data interactions across databases; Developing and operating containerized applications within Kubernetes; Supporting, tuning, and troubleshooting Linux-based systems; Applying core networking knowledge to debug and enhance platform connectivity; Independently diagnosing and resolving complex technical issues; Driving continuous improvement by contributing to CI/CD pipelines
Seniority
Senior, hands-on IC