Senior Mainframe Systems Programmer - Site Reliability Engineering
Core
Senior SRE responsible for reliability, automation, and efficiency of mission-critical z/OS systems, focusing on observability, AI-driven operations, and DevOps integration.
Role type
Senior IC mainframe site reliability engineer
Builds
self-healing workflows, automated incident response playbooks, CI/CD pipelines for COBOL/PL/I applications, and predictive analytics tools
Domain
Mainframe systems (z/OS, CICS, Db2, IMS) and Site Reliability Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
z/OS system programming, JCL, REXX, Python, Infrastructure-as-Code, SRE tenets (SLOs/SLIs, error budgets), AI/ML-driven monitoring, batch job optimization
Preferred skills
AI-driven automation platforms, Zowe Desktop, Dynatrace APM, mainframe open-source ecosystems
Technologies
Ansible, Zowe CLI, z/OSMF, Grafana, Prometheus, IBM Watson AIOps, Splunk ITSI, IBM Dependency-Based Build, UrbanCode Deploy, GitHub, GitLab, Control-M, WLM
Responsibilities
Design and deploy IaC solutions for system provisioning and recovery; Implement predictive analytics tools to detect anomalies; Streamline software delivery pipelines for mainframe applications; Optimize CPU/MIPS utilization and forecast capacity demands; Lead blameless postmortems and reduce MTTR via automated playbooks; Enforce security best practices and develop reusable runbooks
Seniority
Senior, hands-on IC