Engineering Manager, Platform Reliability
Core
Build and lead a new Platform Reliability team to define the reliability roadmap for a rapidly growing global platform, ensuring systems are resilient at scale.
Role type
Engineering Manager, Platform Reliability
Builds
Durable systems, strong operational practices, and high-trust partnerships for a global collaboration platform
Domain
SaaS / Platform Engineering / Distributed Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Engineering management, backend/infrastructure/reliability team leadership, cloud operations (AWS), distributed systems concepts, incident response, SLOs, capacity planning, coaching, technical mentorship
Preferred skills
Curiosity about AI tools, automation over manual labor, long-term solution focus
Technologies
AWS, Kubernetes
Responsibilities
Build and lead a new Platform Reliability team; Partner with technical leaders to define the roadmap for core reliability systems; Establish best-in-class operating practices for incident response and risk management; Work across teams to embed reliability into architecture; Lead end-to-end projects from scoping to rollout; Build a culture of strong engineering execution through coaching and feedback
Seniority
Manager, hands-on IC with people leadership