Staff Site Reliability Engineer - Paze
Core
Define availability standards and implement resiliency patterns in applications and infrastructure to support the U.S. financial system.
Role type
Staff Site Reliability Engineer
Builds
Automation tooling for deployments, configuration changes, disaster recovery, and observability systems.
Domain
Financial services / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Java, Docker, Microservices Architecture, Kafka/SQS/JMS, Oracle/DynamoDB/Aurora, Redis/memcached, Linux administration, CI/CD pipelines (Git/Jenkins), TCP/UDP/IP protocols
Preferred skills
Ruby, JavaScript, Kubernetes, Swarm, 24x7 production environment support
Technologies
Python, Go, Java, Docker, Kafka, SQS, JMS, Oracle, DynamoDB, Aurora, Redis, memcached, Git, Jenkins, Kubernetes, Swarm
Responsibilities
Design and implement software/tools for performance, availability, and scalability; Build automation for application management; Design and evangelize observability systems; Evaluate application capacity and recommend scaling paths; Troubleshoot performance bottlenecks; Serve as technical liaison providing runbooks; Participate in 24x7 on-call rotation; Mentor team members and lead incident response.
Seniority
Staff, hands-on IC with mentorship