Senior Site Reliability Engineer (SRE)
Core
Building, scaling, and maintaining high-scale hybrid infrastructure (on-prem, public cloud, AI/ML Kubernetes) for a performance-driven advertising platform.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
High-availability hybrid infrastructure, internal automation tooling, and IaC pipelines.
Domain
Cloud Infrastructure / DevOps / Advertising Technology
Deliverable
production ML models | infrastructure
Required skills
Linux system internals, network protocols (TCP/IP, DNS, HTTP, gRPC), Infrastructure as Code (Terraform, Ansible, Puppet, ArgoCD, Jenkins), Kubernetes, Docker, Go/Python/Rust programming
Preferred skills
Telemetry/metrics/alerting stacks (Prometheus, Grafana, ELK), cloud cost optimization
Technologies
Kubernetes, Docker, Terraform, Ansible, Puppet, ArgoCD, Jenkins, Fastly, Cloudflare, Akamai, CloudFront, Prometheus, Grafana, ELK, Go, Python, Rust
Responsibilities
Maintain hybrid infrastructure availability and performance, build internal software tooling and manage IaC pipelines, perform deep-dive troubleshooting across the full stack, design and maintain monitoring and alerting setups, lead incident resolution and conduct post-mortems
Seniority
Senior, hands-on IC
