Site Reliability Engineer
Core
Build and maintain distributed, real-time systems serving Point72's Global Macro business through automation, monitoring, and operational optimization.
Role type
Site Reliability Engineer (SRE)
Builds
Distributed real-time systems and foundational SRE program components
Domain
Financial services / High-frequency trading infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python, PowerShell, Linux, Windows, AWS, Terraform, Ansible, Docker, Kubernetes, AWS EKS, AWS ECS, SLO definition, system capacity planning, code review, incident troubleshooting, design review
Preferred skills
C# comprehension, enterprise agile methodology, open-source solution expertise
Technologies
Python, PowerShell, Terraform, Ansible, Docker, Kubernetes, AWS EKS, AWS ECS
Responsibilities
Build foundational technical components for the SRE program; Collaborate with development and quant teams to maintain SLOs; Monitor system capacity and performance to identify bottlenecks; Review and provide feedback on peer automation code; Troubleshoot and resolve system issues; Participate in or lead design reviews for technology and automation strategies
Seniority
Mid-level to Senior, hands-on IC