Manager, Site Reliability Engineering
Core
Lead a team of SREs to ensure the reliability, performance, and scalability of LayerZero's blockchain node infrastructure and platform services.
Role type
Manager, Site Reliability Engineering (SRE)
Builds
Blockchain node infrastructure (validator/full/archive nodes, RPC) and cross-chain dApp platform services
Domain
Blockchain / Distributed Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Team leadership, SRE strategy, Kubernetes at scale, Helm, TypeScript or Golang, Unix/Linux internals, incident response automation, capacity planning
Preferred skills
Deep familiarity with DLTs, architecture decision making
Technologies
Kubernetes, Helm, TypeScript, Golang
Responsibilities
Lead and develop a team of SREs, own reliability strategy for blockchain node infrastructure, partner with Engineering and Product teams, drive infrastructure-as-code practices, establish on-call and incident response processes, stay hands-on with complex incidents
Seniority
Manager, hands-on technical leadership
