Principal SRE (DE & Agentic Tooling)
Core
Architect and evolve the reliability, execution, and validation layers for an AI-embedded fitness platform, ensuring systems are reliable, observable, and production-ready by default.
Role type
Principal Site Reliability Engineer (Platform Engineering)
Builds
Intelligent fitness platform with AI embedded in products, workflows, and decision points
Domain
Fitness industry + Cloud Infrastructure & AI/LLM Systems
Deliverable
production ML models | infrastructure
Required skills
distributed systems design, cloud platforms (AWS), Kubernetes, infrastructure-as-code, reliability engineering practices, CI/CD systems, observability systems, technical standards definition, secure multi-tenant infrastructure
Preferred skills
AI/LLM-powered systems support, high-throughput ephemeral compute systems, internal developer platforms, governance in regulated environments, advanced validation systems (canarying, chaos engineering)
Technologies
AWS, Kubernetes, CI/CD tools, observability stacks (metrics, logging, tracing)
Responsibilities
Architect core platform capabilities for reliability including execution environments and validation pipelines; Design and implement fast, ephemeral, and strictly isolated execution environments; Transform CI/CD into a validation system with automated verification; Build production-like validation environments; Establish deep observability patterns for autonomous workflows; Define and implement guardrails-as-code; Design for reliability including scalability and fault tolerance; Lead technical design reviews; Define and document reusable infrastructure patterns
Seniority
Principal, hands-on IC with technical leadership