Sr. Site Reliability Engineer
Core
Ensure reliability, scalability, and performance of cloud-based systems and applications for a ransomware and breach containment platform.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Highly scalable SaaS service built using cloud-native technologies, deployed on-premises
Domain
Cybersecurity / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
AWS infrastructure management, Azure infrastructure management, Python scripting, Go scripting, PowerShell scripting, CI/CD pipeline design, containerization (Docker, Kubernetes), microservices architecture, root cause analysis, incident response, security compliance
Preferred skills
AWS Solutions Architect certification, Azure DevOps Engineer certification, Azure Security Engineer certification, Kubernetes orchestration
Technologies
AWS, Azure, Docker, Kubernetes, Azure DevOps, Jenkins, GitLab CI/CD, Python, Go, PowerShell
Responsibilities
Monitor system performance, application health, and infrastructure metrics; perform oncall duty for production uptime and customer escalations; execute release upgrades, maintenance activities, and hotfixes; lead incident response and resolution efforts; implement security best practices and controls; drive continuous improvement initiatives for infrastructure and services
Seniority
Senior, hands-on IC