Principal Site Reliability Engineer
Core
Design, build, and maintain shared infrastructure services and platforms (Kubernetes, CI/CD, observability) for product and application teams in a multi-cloud environment.
Role type
Principal Site Reliability Engineer (Platform Engineering)
Builds
Reusable, reliable, and scalable shared infrastructure services and platforms
Domain
SaaS, Multi-cloud (AWS, Azure, GCP), Distributed Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes (single/multi-tenant), Go (Golang), Python, Event-Driven Architecture, CI/CD pipelines, Distributed Systems design, Observability, RESTful API design, Service Mesh, Relational Databases
Preferred skills
Multi-cloud abstractions, Kafka/Pub/Sub, GitLab CI/ArgoCD, Prometheus/Grafana/ELK/Datadog, Istio/Envoy, MySQL/PostgreSQL
Technologies
Kubernetes, Go, Python, AWS, Azure, GCP, Kafka, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK, Datadog, Istio, Envoy, MySQL, PostgreSQL
Responsibilities
Architect and manage highly available Kubernetes platforms; Develop internal tools and automation for infrastructure provisioning; Design and implement shared Event-Driven Architecture components; Build resilient Distributed Systems components; Manage and optimize shared infrastructure across Multi-Region Cloud Environments; Establish centralized Observability and Monitoring platforms; Define and implement RESTful API designs for infrastructure services; Implement and manage Service Mesh capabilities; Design and optimize Relational Database services; Participate in on-call rotations for critical shared infrastructure
Seniority
Principal, hands-on IC with architectural influence
