Principal Site Reliability Engineer (we have office locations in Cambridge, Leeds and London)
Core
Establish and grow an organization-wide Site Reliability Engineering (SRE) capability, leading a small team to improve reliability across platforms and embed SRE principles into engineering squads.
Role type
Principal Site Reliability Engineer (Engineering Leader)
Builds
SRE team, reliability standards, tooling, and services that reduce operational toil
Domain
Healthcare genomics, bioinformatics pipelines, cloud infrastructure
Deliverable
production ML models | infrastructure | product features
Required skills
SRE principles and practices, software engineering (Python), system architecture resilience, platform engineering (CI/CD, IaC, monitoring), cloud platform experience, team leadership and mentorship
Preferred skills
Healthcare or bioinformatics background, regulated environment experience
Technologies
AWS, Python, NextFlow, Prefect, Dremio, Terraform, GitLab, Artifactory, DataDog, ECS, Fargate, Kubernetes, HPC
Responsibilities
Build and lead a small SRE team, identify high-value reliability opportunities, define standards and tooling, develop services to reduce toil, partner with product teams to embed SRE practices, influence engineering strategy and governance
Seniority
Principal, hands-on engineering leader with strategic scope