Site Reliability Engineer
Core
Scaling PolarGrid's edge inference network from a small cluster to a geographically distributed system, owning the full lifecycle of node provisioning and readiness.
Role type
Senior SRE (distributed GPU inference systems)
Builds
A global, low-latency inference network for LLMs and voice pipelines (STT, TTS) across 20+ locations
Domain
Cloud infrastructure, distributed systems, GPU computing, real-time AI inference
Deliverable
production ML models | infrastructure
Required skills
Distributed systems operations, Kubernetes production management, Networking fundamentals (routing, TCP/UDP, load balancing), GPU workload management, Observability (metrics, tracing, latency analysis), Infrastructure as Code (Terraform, AWS CDK), Automation and scripting, Capacity planning, Failure handling and resilience design
Preferred skills
Helm-based GitOps, Triton Inference Server multipod deployments, Real-time voice/WebRTC stacks (LiveKit), CDN or edge compute background, ACME certificate automation
Technologies
Kubernetes, Helm, Triton, CUDA, AWS CDK, Terraform, LiveKit, S3, AWS
Responsibilities
Automate cluster bring-up and node lifecycle management, Operate and optimize the global routing and inference layer for TTFT and tail latency, Ensure model serving reliability across LLM and voice pods, Build observability and resilience systems for the distributed GPU network, Develop internal tooling and deployment workflows for safe global rollouts
Seniority
Senior, hands-on IC