Senior Site Reliability Engineer
Core
Own and evolve the systems responsible for serving millions of machine learning inference requests daily, ensuring reliability, performance, and scalability for GPU-based AI infrastructure.
Role type
Senior Site Reliability Engineer (Machine Learning Infrastructure)
Builds
Cloud-agnostic infrastructure solutions for ML inference workloads, including load balancing, autoscaling, queuing, and workload orchestration.
Domain
E-commerce, Machine Learning Infrastructure, Cloud Systems
Deliverable
production ML models
Required skills
Designing and operating large-scale distributed systems, load balancing, autoscaling, queuing systems, traffic management, low-latency real-time backend systems, building resilient redundant infrastructure, deploying containerised workloads at high scale, supporting variable-duration workloads, pragmatic technical decision-making, cross-team collaboration, ownership.
Preferred skills
Experience supporting machine learning infrastructure, GPU workloads.
Technologies
Datadog, GPU-based systems, containerised workloads
Responsibilities
Design and build cloud-agnostic infrastructure solutions, work across the full infrastructure lifecycle from architecture to incident management, build systems for load balancing and autoscaling, monitor production systems and analyse usage patterns, partner with ML engineers to ensure service reliability and cost-efficiency, identify bottlenecks and improve deployment workflows.
Seniority
Senior, hands-on IC