Senior / Staff Site Reliability, Platform Engineering
Core
Design, build, and maintain shared infrastructure services and platforms (Kubernetes, Cloud, Event-Driven) to ensure reliability and scalability for product teams.
Role type
Staff Platform Engineer (SRE)
Builds
Shared infrastructure services, Kubernetes platforms, CI/CD pipelines, Event-Driven components, Observability tools, and Distributed Systems building blocks.
Domain
Cloud-native SaaS, Identity Security, Multi-cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes (production, multi-tenant), Go (Golang), Python, Cloud Providers (AWS/GCP/Azure), Event-Driven Architecture (Kafka/RMQ/NATS), CI/CD (GitLab CI/ArgoCD), Distributed Systems, Observability (Prometheus/Grafana/ELK/Datadog), RESTful API design, Service Mesh (Istio/Envoy), Relational Databases (MySQL/PostgreSQL)
Preferred skills
Multi-cloud abstractions, Multi-Region Cloud Environments
Technologies
Kubernetes, Go, Python, AWS, Azure, GCP, Kafka, GitLab CI, ArgoCD, Prometheus, Grafana, ELK, Datadog, Istio, Envoy, MySQL, PostgreSQL
Responsibilities
Architect and manage highly available Kubernetes platforms as a service; Develop internal tools and automation for infrastructure provisioning; Design and implement shared Event-Driven Architecture components; Build resilient Distributed Systems components; Manage and optimize shared infrastructure across Multi-Region Cloud Environments; Establish centralized Observability and Monitoring platforms; Define and implement RESTful API designs for infrastructure services; Implement and manage Service Mesh capabilities; Design and optimize Relational Database services; Collaborate with product teams on infrastructure needs; Participate in on-call rotations.
Seniority
Staff, hands-on IC with technical leadership

