Principal Site Reliability Engineer - Paze
Core
Apply software and systems engineering practices to improve the reliability, resilience, scalability, and operational health of production services at an enterprise scope.
Role type
Principal Site Reliability Engineer
Builds
Production services, CI/CD pipelines, observability tooling, and automation for financial transaction systems
Domain
Financial services / Cloud Infrastructure / Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Software engineering, distributed systems, automation, observability, public cloud (AWS), Linux/Unix, incident response, capacity management, SLIs/SLOs, Infrastructure as Code
Preferred skills
High availability system design, container orchestration, disaster recovery, reusable platform development, technical leadership
Technologies
AWS, Azure, GCP, OCI, Linux, Kubernetes, Docker, Terraform, Prometheus, Grafana, ELK
Responsibilities
Define and implement SLIs, SLOs, and error budgets; lead enterprise-level incident response and post-mortems; drive continuous improvement in CI/CD and deployment practices; reduce operational toil through automation; partner with engineering teams to embed reliability into the development lifecycle; establish enterprise technical direction for reliability and resilience.
Seniority
Principal, enterprise-level technical leadership and strategy
