Senior App/Prod Support (Tier 3 Site Reliability Engineer (SRE) / Platform Engineer)
Core
Own platform reliability practices, drive DevOps automation, and provide Tier 3 troubleshooting for high-complexity incidents in a cloud-native environment.
Role type
Senior IC Site Reliability Engineer (SRE) / Platform Engineer
Builds
Cloud infrastructure, CI/CD pipelines, observability stacks, and automation tooling for microservices.
Domain
Telecom / Cloud Infrastructure / Streaming Data
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (AKS), CI/CD engineering (GitHub Actions), Python scripting, Observability (Prometheus/Grafana/AlertManager), Streaming stack (Confluent Kafka, Azure Event Hub, Apache Flink), Incident leadership, Governance controls
Preferred skills
Postgres performance operations, Telecom-scale HA systems experience
Technologies
Kubernetes, Azure Kubernetes Service (AKS), GitHub Actions, JFROG Helm, Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, Airflow, Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, Apache Flink
Responsibilities
Own platform reliability practices for availability, resilience, and latency; Drive DevOps and automation initiatives including Golden Image improvements; Implement and maintain CI/CD pipelines; Lead cloud infrastructure creation, maintenance, and governance; Provide troubleshooting support for high-complexity incidents; Lead capacity planning and disaster recovery planning
Seniority
Senior to Lead IC (10-17 years experience)