Service Reliability Engineer
Core
Administer and optimize a large-scale Kubernetes platform (hundreds of clusters, thousands of nodes) to ensure performance, reliability, and cost-effectiveness for trading desks and data scientists.
Role type
Senior IC Service Reliability Engineer (Kubernetes/Container Infrastructure)
Builds
Scalable on-premises and cloud-based container orchestration services
Domain
Financial Trading / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes cluster administration, networking fundamentals, storage fundamentals, containerization fundamentals, AWS or GCP familiarity, automation scripting, cross-team collaboration
Preferred skills
Kubernetes operators (Flux, Fluentbit, Fluentd, Prometheus), on-premise infrastructure administration, full stack web application development, multi-language programming
Technologies
Kubernetes, AWS, Google Cloud Platform, Flux, Fluentbit, Fluentd, Prometheus
Responsibilities
Ensure successful onboarding and deployment of applications into the Kubernetes platform, relentlessly automate repetitive manual tasks, solve complex problems to build new features/components, engage with internal customers to ensure platform reliability and cost optimization
Seniority
Senior, hands-on IC