Senior Software Development Engineer (Site Reliability)
Core
Ensuring the reliability, availability, performance, and operational scalability of the myPBM platform through automation, observability, and incident management.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Stable, scalable delivery of client-facing services for the myPBM platform
Domain
Healthcare technology / Cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Site reliability engineering, DevOps, platform engineering, Monitoring and observability tools (Splunk, AppDynamics), Cloud platforms (Azure, AKS, Kubernetes), CI/CD pipelines (GitHub Actions, Jenkins), Incident management, Root cause analysis, Infrastructure and networking fundamentals, Scripting (Python, Bash, PowerShell)
Preferred skills
Healthcare or regulated environment experience, SRE principles (SLIs, SLOs, error budgets), DevSecOps practices, Large-scale distributed systems support
Technologies
Azure, AKS, Kubernetes, GitHub Actions, Jenkins, Splunk, AppDynamics, xMatters, MIR3, Python, Bash, PowerShell
Responsibilities
Define and manage SLIs, SLOs, and SLAs; Lead incident response and root cause analysis; Implement end-to-end observability; Support and enhance CI/CD pipelines; Manage cloud infrastructure and disaster recovery; Implement continuous security monitoring and vulnerability remediation; Automate operational tasks using infrastructure as code
Seniority
Senior, hands-on IC