CareerPlanGet AI match score →

Senior Site Reliability Engineer

💼 Full-time🗓 2026-06-25

Core

Design and build robust, scalable, and fault-tolerant infrastructure and services for a high-throughput healthcare platform.

Role type

Senior Site Reliability Engineer (IC)

Builds

AWS-based platform, self-healing systems, CI/CD pipelines, observability tooling

Domain

Healthcare / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

AWS, Kubernetes, distributed system architecture, infrastructure as code (Terraform), observability (Grafana, OpenTelemetry, Prometheus), incident response, CI/CD pipeline design, chaos engineering, disaster recovery

Preferred skills

Java, Python, Go, GitOps best practices, mentorship

Technologies

AWS, Kubernetes, Terraform, Grafana, OpenTelemetry, Prometheus, Datadog, Git

Responsibilities

Design and implement scalable, fault-tolerant infrastructure and services on AWS and Kubernetes; Define and help drive adoption of SLIs, SLOs, and SLAs; Own and improve observability using Grafana, OpenTelemetry, and related tooling; Build and maintain infrastructure as code (Terraform) and contribute to GitOps best practices; Participate in incident response and on-call rotation; Support the growth of junior and mid-level SRE engineers through mentorship, code reviews, and knowledge sharing; Drive continuous improvements in CI/CD pipelines, service ownership, chaos engineering, disaster recovery, and secure deployments.

Seniority

Senior, hands-on IC

Rewrite
## About the role We are looking for a Senior Site Reliability Engineer to join our Infrastructure Engineering team in Kraków. In this role, you will be a key technical contributor driving the reliability, scalability, and performance of our healthcare platform. You will design and build robust systems, develop automation and tooling, and apply strong software engineering principles to infrastructure challenges at scale. Working within a collaborative team, you will contribute to shaping the architecture of our AWS-based platform, own meaningful portions of our observability and deployment strategies, and help maintain a culture of engineering excellence and continuous improvement. ## Responsibilities - Design and implement scalable, fault-tolerant infrastructure and services on AWS and Kubernetes, contributing to the architecture of self-healing systems. - Collaborate with Product, Engineering, and Security teams to align SRE work with platform and business priorities. - Define and help drive adoption of SLIs, SLOs, and SLAs to maintain consistent performance and high reliability across the platform. - Own and improve observability using Grafana, OpenTelemetry, and related tooling to provide deep visibility into system health. - Build and maintain infrastructure as code (Terraform) and contribute to GitOps best practices across the team. - Participate in incident response and on-call rotation, including writing postmortems and driving long-term remediation efforts. - Support the growth of junior and mid-level SRE engineers through mentorship, code reviews, and knowledge sharing. - Contribute to the reliability roadmap for high-throughput, real-time systems in healthcare operations. - Drive continuous improvements in CI/CD pipelines, service ownership, chaos engineering, disaster recovery, and secure deployments. ## What You Bring - Experience in Site Reliability Engineering, Cloud Infrastructure, or Platform Engineering. - Software engineering experience building production-grade systems (Java, Python, Go, or similar). - Solid track record working on high-traffic, mission-critical platforms in SaaS, IoT, or healthcare environments. - Strong expertise in cloud platforms (especially AWS), Kubernetes, and distributed system architecture. - Hands-on experience with monitoring, logging, and observability tools (Prometheus, OpenTelemetry, Datadog, etc.). - Good knowledge
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗