Site Reliability Engineer (SRE)
Core
Ensure stability and high performance of critical systems across cloud, mainframe, and legacy environments through monitoring, automation, and resilience engineering.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Automated monitoring, alerting, self-healing workflows, and chaos engineering experiments for financial services clients.
Domain
Financial services, cloud infrastructure, and legacy mainframe systems.
Required skills
Linux/Unix systems administration, cloud platform management (AWS, Azure, OpenShift), containerization and orchestration (Kubernetes, Docker), incident management, chaos engineering, infrastructure as code, networking, database systems, event streaming (Kafka), service mesh (Istio), job scheduling (CA7), middleware technologies.
Preferred skills
Experience in financial services or regulated industries, relevant certifications (AWS/Azure architecture, RHCE, VCP, CKA/CKAD).
Technologies
AWS Lambda, Azure Cloud, Confluence, Datadog, DevOps, Docker, Dynatrace, Grafana, Istio, ITSM, JIRA, Kafka, Kubernetes, Linux, OpenShift, Oracle, Phoenix, Prometheus, RHEL, SQL, ServiceNow, Splunk, Unix, Windows, VMware
Responsibilities
Coordinate responses to critical events with application support teams; triage and respond to alerts from event correlation platforms; participate in on-call rotations for 24/7 coverage; conduct blameless post-mortems and root cause analyses; design and implement automated monitoring and alerting systems; create robust dashboards and establish SLAs/SLOs; analyze metrics for performance tuning and fault detection; develop and implement chaos engineering practices; design fault injection experiments to validate system resilience; build self-healing capabilities and automated remediation workflows; manage infrastructure across mainframe, Windows, RHEL, and cloud platforms; maintain virtualization and storage systems; utilize ITSM tools for incident management and issue tracking; identify opportunities to enhance application stability; maintain comprehensive knowledge bases and runbooks; mentor junior team members on resiliency patterns.
Seniority
Senior, hands-on IC

