Principal Site Reliability Engineer - SaaS
Core
Ensuring the availability, reliability, and performance of mission-critical financial systems and services.
Role type
Principal Site Reliability Engineer (IC6)
Builds
Cloud-based infrastructure, monitoring systems, and automated operational processes for investment management solutions.
Domain
FinTech / Investment Management
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering principles, Cloud platform expertise (AWS, Azure, GCP), Infrastructure as Code (IaC), Automation & Scripting, Monitoring & Observability, Incident Management & Root Cause Analysis, Capacity Planning, Mentorship
Preferred skills
Deep knowledge of SRE practices, Experience with emerging technologies, Data-driven analytical mindset
Technologies
AWS, Azure, GCP, IaC tools, Monitoring/Alerting/Logging systems
Responsibilities
Lead design and implementation of systems to improve reliability and scalability; Automate operational tasks and processes; Monitor system performance and proactively resolve issues; Collaborate with development teams to integrate reliability practices; Conduct root cause analysis for incidents; Develop and maintain infrastructure monitoring and logging systems; Provide mentorship to junior engineers; Engage in capacity planning and performance tuning.
Seniority
Principal, hands-on IC with mentorship responsibilities