Senior Cluster Site Reliability Engineer
Core
Scale research compute clusters to support AI/ML research in finance, ensuring high uptime and reliability for on-prem and cloud infrastructure.
Role type
Senior Cluster Site Reliability Engineer (HPC/ML Infrastructure)
Builds
High-performance computing (HPC) platforms and batch compute systems for machine learning research
Domain
Finance, Artificial Intelligence, Machine Learning, High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
SRE/DevOps, HPC/batch compute frameworks, Kubernetes, Infrastructure-as-Code, Cloud infrastructure, Observability, Distributed storage, Scripting
Preferred skills
HPC frameworks (Slurm, Grid Engine), Kubernetes job orchestrators, Distributed computing frameworks, ML frameworks, Containerization, HPC networking, Security/IAM
Technologies
Slurm, Kueue, AWS Batch, GCP Batch, Kubeflow, MLflow, Horovod, Terraform, Ansible, AWS, GCP, Prometheus, Grafana, Loki, ELK, OpenTelemetry, Lustre, Ceph, S3, Docker, Podman, Singularity, InfiniBand, RDMA
Responsibilities
Triage and resolve urgent cluster outages; Define and track SLAs for cluster uptime; Diagnose systemic issues and engineer precision solutions; Develop custom observability mechanisms; Design policies for fair cluster usage; Forecast cluster growth and optimize cost/usability
Seniority
Senior, hands-on IC