Staff, Site Reliability Engineer(Global Security)
Core
Senior SRE leading reliability, automation, and incident response for Identity and Access Management (IAM) systems in hybrid-cloud environments.
Role type
Staff Site Reliability Engineer (IAM)
Builds
Resilient, highly available IAM infrastructure, self-service automation platforms, and CI/CD pipelines.
Domain
Financial Services / Identity and Access Management (IAM) / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE/DevOps leadership, software engineering (Python/Go/Java), hybrid/multi-cloud architecture, CI/CD ownership, container orchestration (Kubernetes), observability (Prometheus/Grafana/SIEM), incident management, disaster recovery planning.
Preferred skills
IAM platform experience (Okta/Microsoft Entra/HashiCorp Vault), Infrastructure as Code (Terraform/Ansible), identity protocols (OAuth2/OIDC/SAML), AIOps.
Technologies
Kubernetes, Terraform, Ansible, Python, Go, Java, Jenkins, GitLab CI, GitHub Actions, Prometheus, Grafana, Splunk, SIEM, AWS, Azure, Docker, Helm, PowerShell, Bash.
Responsibilities
Define and maintain SLOs/SLIs/error budgets; design resilient IAM infrastructure across multi-region/hybrid-cloud; write production-grade code for services and automation; lead high-severity incident response and postmortems; build self-service platforms and golden-path automation; orchestrate CI/CD pipelines and release engineering; develop failover strategies and conduct chaos engineering; mentor engineers on reliability practices.
Seniority
Staff, hands-on IC with strategic leadership