Principal Site Reliability Engineer
Core
Designing, building, and maintaining shared infrastructure services and platforms for product and application teams in a multi-cloud environment.
Role type
Principal Site Reliability Engineer (Platform Engineering)
Builds
Reusable, reliable, and scalable Kubernetes platforms, event-driven architecture components, CI/CD pipelines, and distributed system building blocks.
Domain
Cloud Infrastructure & Platform Engineering
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes (production, multi-tenant), Go (Golang), Python, Cloud Providers (AWS/Azure/GCP), Event-Driven Architecture, CI/CD pipelines, Distributed Systems, Observability, RESTful API design, Service Mesh, Relational Databases
Preferred skills
Multi-cloud abstraction, Kafka/Pub/Sub, GitLab CI/ArgoCD, Prometheus/Grafana/ELK/Datadog, Istio/Envoy, MySQL/PostgreSQL
Technologies
Kubernetes, Go, Python, AWS, Azure, GCP, Kafka, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK, Datadog, Istio, Envoy, MySQL, PostgreSQL
Responsibilities
Architect and manage highly available Kubernetes platforms as a service; Develop internal tools and automation for infrastructure provisioning; Design and optimize foundational cloud solutions and reusable patterns; Implement shared Event-Driven Architecture and messaging platforms; Build and maintain CI/CD pipelines as a service; Design resilient distributed system components; Manage multi-region cloud environments; Establish centralized observability and monitoring platforms; Define and implement RESTful API designs for infrastructure services; Manage shared relational database services; Collaborate with product teams on infrastructure needs; Participate in on-call rotations.
Seniority
Principal, hands-on IC with architectural influence
