Principal, SRE
Core
Define architectural vision and resiliency strategy for critical platforms, embedding reliability principles into system design decisions and lifecycle governance at enterprise scale.
Role type
Principal Site Reliability Architect / Engineer
Builds
Resiliency guidelines, reference architectures, automation logic for infrastructure configuration and application lifecycle management, and observability capabilities.
Domain
Financial services / Cloud and on-prem infrastructure / Distributed systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Resiliency architecture design, distributed systems knowledge, observability standards (metrics, logging, tracing), Infrastructure as Code, CI/CD pipeline automation, scripting/programming (Python, C#, Java), IaC/Configuration Management (Terraform, Ansible, etc.), incident root cause analysis
Preferred skills
Enterprise-scale system architecture, maturity model development, community of practice leadership
Technologies
Terraform, Bicep, ARM, Chef, Puppet, Ansible, Python, C#, Java
Responsibilities
Create and maintain resiliency guidelines; guide adoption of resilient software and infrastructure architectures; provide guidance in incident response and root cause analysis; build infrastructure configuration and application management automation; foster a broader community of practice
Seniority
Principal, hands-on IC with strategic influence