Lead Site Reliability Engineer
Core
Lead SRE responsible for designing, implementing, and maintaining highly available, scalable, and resilient systems across cloud infrastructure for the Federal Reserve Bank of San Francisco.
Role type
Lead Site Reliability Engineer (IC with mentorship) (via careerplan.io/jobs/R-0000033169-1-lead-site-reliability-engineer-at-rb)
Builds
Cloud infrastructure, deployment pipelines, monitoring solutions, and internal tools for critical financial systems
Domain
Financial services / Cloud Infrastructure / DevOps
Required skills
Java, Python, Node.js, microservices architecture, distributed systems, AWS services (Lambda, ECS, EC2, Fargate, S3, RDS, DynamoDB, Aurora, VPC, Route53, CloudFront, API Gateway), Terraform, GitLab CI/CD, Docker, Kubernetes, SAST/DAST, Grafana, Datadog, Splunk, X-Ray
Preferred skills
GenAI/LLM applications, serverless architectures, event-driven systems, chaos engineering, multi-cloud environments, AWS certifications (Solutions Architect, DevOps Engineer)
Technologies
Java, Python, Node.js, AWS, Terraform, GitLab, Docker, Kubernetes, Grafana, Datadog, Splunk, X-Ray
Responsibilities
Design and maintain resilient cloud systems; establish and monitor SLIs/SLOs/SLAs; lead incident response and root cause analysis; architect and manage AWS infrastructure using IaC; automate deployment pipelines and operational workflows; mentor junior SREs and promote SRE culture; conduct code reviews and provide technical guidance; integrate security practices into CI/CD pipelines; document processes and runbooks
Seniority
Senior, hands-on IC with leadership responsibilities