Director PCS Cloud Operations SRE
Required skills
Bachelor’s degree in computer science or a STEM field, 10 years experience in leading technical teams in complex, fast-paced environments, 5+ years of in Cloud Ops and SRE leadership roles, Proven expertise in the areas of DevSecOps, Day-2 Ops, APM/RUM, and Cloud Operations, Proficiency building and operating services on public cloud (AWS-first) with CI/CD and Infrastructure-as-Code (e.g., Terraform/CloudFormation), Track record establishing SLIs/SLOs/SLAs, observability, and incident/change management at scale, Strong leadership and team management skills, with the ability to inspire and motivate a team of engineers, Excellent project management skills, with the ability to manage multiple complex projects simultaneously, In-depth knowledge of SaaS technologies, cloud computing, and medical device development processes
Preferred skills
Experience scaling CloudOps/SRE for multiple products and customer deployments, Deep fluency in SLI/SLO/SLA design, error budgets, runbooks, and auto-healing patterns, Strong AWS architecture and operations; Well-Architected reviews; capacity and cost optimization (FinOps), Modern observability (APM/RUM/logs/metrics/traces) and AIOps for predictive analytics/anomaly detection, Security by design (DevSecOps, policy-as-code) and DR/BCP planning/testing
Technologies
AWS, Terraform, CloudFormation, CI/CD, Infrastructure-as-Code, APM/RUM, logs, metrics, traces, observability, FinOps, DevSecOps, policy-as-code, SLI/SLO/SLA, error budgets, runbooks, auto-healing, AIOps, DR/BCP, Well-Architected, cost allocation, right-sizing, savings plans/reserved instances, spend governance, unit-economics optimization
Responsibilities
Serve as the functional leader for the PCS Digital Cloud Operations team. Define the operating model, governance, and KPIs; drive automation and observability; and ensure secure, reliable deployments across environments with continuous improvement and tight collaboration with security. This role reports to the VP of Engineering – PCS Apps & Platform. Key responsibilities include: Own Cloud Operations for PCS cloud applications; stand up and scale CloudOps capabilities to support multiple products while adhering to committed SLAs. Institutionalize SRE practices: implement SLI/SLO/SLA frameworks, error budgets, incident/post-mortem processes, and reliability runbooks; champion automation to reduce toil and improve service health and monitoring. Build end-to-end observability (APM/RUM, logs, metrics, traces, health dashboards, proactive alerting) and evolve toward auto-healing and AIOps for anomaly detection and closed-loop remediation. Drive change, incident, and problem management with clear RACI and stakeholder communications; reduce MTTR through streamlined L1–L4 escalation. Establish and test DR/BCP posture; conduct AWS Well-Architected and operational readiness reviews for services (AWS-first, with multi-cloud considerations as needed). Lead FinOps practices: cost allocation and accountability, right-sizing, savings plans/reserved instances, spend governance, and unit-economics optimization. Evolve the operating model in partnership with platform and application teams; standardize CI/CD templates and "everything-as-code" for speed and repeatability. Build and develop a high-performing team: hire, coach, and grow CloudOps/SRE talent and the next set of leaders; uphold high standards for quality and customer satisfaction
Seniority
Director
Domain
Cloud Operations, SRE, DevSecOps, FinOps, APM/RUM, Observability, Cloud Computing, SaaS, Medical Device Development