Platform ULL - Colo - Reliability
Core
Ultra Low Latency Platform Engineer managing global colocation infrastructure, ensuring high-performance compute environments for trading venues.
Role type
Senior IC infrastructure reliability engineer (low latency)
Builds
Global colocation estate with 400+ servers across 30 sites, petabyte-scale storage systems, and automated recovery processes.
Domain
Financial technology infrastructure, low latency trading systems
Deliverable
production ML models | product features | infrastructure
Required skills
Linux operations (RHEL/CentOS/Rocky), server hardware management (HP/SuperMicro/Dell), kernel bypass optimization (Solarflare/Mellanox), configuration management (Chef/Ansible), observability (Grafana/Prometheus), scripting (Python/Ruby/Bash), network stack tuning (TCP/UDP/NTP/PTP), root cause analysis, capacity management, incident response.
Preferred skills
Experience with trading venues (Nasdaq/LSE/Euronext), overclocked server management.
Technologies
RHEL, Solarflare, Mellanox, Chef, Ansible, Grafana, Prometheus, Python, Ruby, Bash, Wireshark, Tshark
Responsibilities
Manage distributed compute environment and multiple petabyte-scale storage systems; Install, manage, and monitor Linux operating systems; Troubleshoot complex hardware and software issues; Create self-healing systems and automated recovery processes; Respond to system incidents and participate in on-call rotations; Conduct root cause analysis of incidents and outages; Reduce operational toil through development of automated workflows.
Seniority
Senior, hands-on IC