Senior Cluster Site Reliability Engineer
Core
Scale research compute clusters to support AI/ML research in finance, ensuring high uptime and reliability for technical staff.
Role type
Senior Site Reliability Engineer (HPC/Cluster Operations)
Builds
World-class HPC platform for researchers to solve cutting-edge machine learning problems at scale
Domain
Finance / High-Performance Computing / Machine Learning Infrastructure
Deliverable
infrastructure
Required skills
SRE/DevOps, HPC/batch compute frameworks (Slurm, Kueue), cloud infrastructure (AWS/GCP), infrastructure-as-code (Terraform, Ansible), observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry), distributed storage (Lustre, Ceph, S3), scripting (Python, Ruby)
Preferred skills
Kubernetes-based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), distributed computing frameworks (Ray, Modin, Dask, Spark), ML frameworks (PyTorch, Tensorflow, JAX, Horovod, DeepSpeed), containerization (Docker, Podman, Singularity), HPC networking (InfiniBand, RDMA), security/IAM foundations
Technologies
Slurm, Kueue, AWS, GCP, Terraform, Ansible, Prometheus, Grafana, Loki, ELK, OpenTelemetry, Lustre, Ceph, S3, Python, Ruby, Kubernetes, Airflow, Kubeflow Pipelines, Ray, Modin, Dask, Spark, PyTorch, Tensorflow, JAX, Horovod, DeepSpeed, Docker, Podman, Singularity, InfiniBand, RDMA
Responsibilities
Be a first responder for cluster outages and triage urgent issues; ensure high cluster uptime and define/track SLAs; diagnose systemic problems and engineer precision solutions; develop robust metrics and custom observability mechanisms; help design policies and enforcement mechanisms for fair cluster usage; assist in forecasting cluster growth and optimizing operations for cost and usability
Seniority
Senior, hands-on IC