Platform Reliability Engineer
Core
Manage and optimize server and network infrastructure for a large financial buy-side organization to ensure production reliability and reduce operational overhead.
Role type
Senior Platform Reliability Engineer (SRE)
Builds
High-performance Linux-based platform for systematic financial strategies
Domain
Financial services / High-Performance Computing (HPC)
Deliverable
production ML models | infrastructure
Required skills
Linux system internals, storage technologies, IT infrastructure components, system automation, container orchestration, HPC infrastructure, AI/ML for infrastructure
Preferred skills
GPFS, Salt, Kubernetes, Nomad, VMware, Slurm, GPU, AI/ML implementation
Technologies
Linux, GPFS, Salt, Kubernetes, Nomad, VMware, Slurm, GPU
Responsibilities
Ensure production reliability of Linux-based platform, provide rapid emergency response to infrastructure issues, develop and enhance observability platform, identify risks and implement contingency plans, participate in on-call rotations, contribute to organizational knowledge through documentation
Seniority
Senior, hands-on IC