Principal Site Reliability Engineer
Core
Partner with software development teams to build reliable, scalable, secure, and cloud-native services for a global fitness platform serving 30K clubs and 40M members.
Role type
Principal Site Reliability Engineer (IC)
Builds
Cloud-native services, deployment pipelines, and operational guardrails for a SaaS fitness management platform.
Domain
SaaS / Fitness Technology / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Site Reliability Engineering, DevOps, cloud infrastructure, software engineering, architecture influence, incident response, root cause analysis, mentoring, programming (Go/Python/Node.js/Java), infrastructure as code, CI/CD, observability, database management, AI-assisted engineering tools.
Preferred skills
Agentic services/AI-enabled tools, SaaS service selection, developer productivity tooling, security-minded engineering practices.
Technologies
AWS (ECS, Fargate, Lambda), Terraform, GitHub Actions, CircleCI, Jenkins, Honeycomb.io, New Relic, CloudWatch, Grafana, Aurora, MySQL, PostgreSQL, MongoDB, DynamoDB, Redshift, SQL Server.
Responsibilities
Define and improve Service Level Indicators/Objectives, establish best practices for availability and reliability, design deployment pipelines, lead production incident response and post-incident learning, mentor engineers, integrate security into SDLC.
Seniority
Principal, hands-on IC with strategic influence