Site Reliability Champion, Specialist
Core
Provides subject matter expertise and coordination for site reliability efforts, ensuring system reliability through SLOs, automation, and incident management.
Role type
Senior Site Reliability Champion / Specialist
Builds
Scalable, reliable systems and automated operational workflows for Vanguard's investment management platform.
Domain
Financial Services / Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Incident management, system performance optimization, automation of operational tasks, reliability engineering principles, cross-functional collaboration, chaos experimentation, root cause analysis, KPI definition and tracking, scalable system design
Preferred skills
Master's degree in CS/IT/SE, experience in development or operational support functions
Technologies
AWS (EC2, ECS, Lambda, S3, CloudWatch), Java, Node.js, Angular, Docker, Git, GitHub, JIRA, Confluence, Splunk, Honeycomb, Cucumber, JUnit, Playwright, Puppeteer
Responsibilities
Coordinate cross-product chaos experimentation; maintain centralized incident response playbooks; facilitate blameless post-incident reviews; communicate and enforce reliability standards; drive automation of routine operational tasks; lead incident response efforts; define and track system performance metrics; collaborate with development teams on architecture alignment
Seniority
Senior, hands-on IC with coordination responsibilities