Senior Principal Infrastructure Services (SRE Practice)
Core
Designing and deploying observability services, automation tools, and reliable distributed systems to ensure the performance and resilience of Northern Trust's global IT landscape.
Role type
Senior Principal Site Reliability Engineer (SRE)
Builds
Production observability platforms, automated operational workflows, and highly reliable distributed systems
Domain
Financial services / Cloud infrastructure / Distributed systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Distributed systems architecture, Observability (metrics/logs/traces), Automation & scripting, Incident management & root cause analysis, Capacity planning, CI/CD integration, Container orchestration, Infrastructure as Code (IaC), Technical leadership, System design
Preferred skills
Chaos engineering, Load testing, Mentoring technical teams, Agile/DevOps practices
Technologies
Python, Go, Java, Ruby, Kubernetes, Docker, Prometheus, Grafana, ELK Stack, Terraform, Jenkins, GitLab CI
Responsibilities
Lead design of scalable distributed systems, Drive automation-first approach to reduce operational toil, Architect end-to-end observability solutions, Lead incident response and post-mortem analysis, Define and manage SLIs/SLOs and error budgets, Mentor and develop high-performing technical teams
Seniority
Principal, hands-on IC with strategic influence