Lead Site Reliability Engineer | Production Infrastructure
Core
Lead Site Reliability Engineer managing and mentoring engineers while contributing hands-on to Jump's global production trading environment.
Role type
Lead Site Reliability Engineer (Production Infrastructure)
Builds
High-performance monitoring/alerting systems, real-time packet/flow analysis tooling, automation frameworks, and production tooling.
Domain
Financial technology / Global trading infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Leadership across distributed teams, solving reliability challenges in large-scale production, strategic thinking, Python, Go
Preferred skills
None stated
Technologies
Python, Go
Responsibilities
Architect and implement high-performance monitoring and alerting systems; Oversee and improve incident, change, and post-incident review processes; Identify and eliminate operational toil through automation; Partner with engineering, networking, and trading teams globally; Investigate low-level performance issues across complex software stacks; Influence strategic direction of production tooling and infrastructure scaling.
Seniority
Lead, hands-on IC with mentorship