Principal Systems Engineer
Core
Ensure reliability, scalability, and performance of hosted healthcare platforms through software and systems engineering.
Role type
Senior IC Site Reliability Engineer (SRE)
Builds
Cloud and hybrid healthcare service environments
Domain
Healthcare technology / Cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident management, Root cause analysis (RCA), Service Level Objectives (SLOs) definition, Proactive monitoring, Infrastructure automation, Cloud environment administration, Database performance tuning, Windows Server administration, SQL Server optimization, Networking fundamentals, ITSM processes
Preferred skills
PowerShell scripting, Python scripting, Infrastructure as Code (Terraform, ARM, Bicep), CI/CD pipeline design, Kubernetes management, SLO/SLI implementation, Healthcare domain experience
Technologies
LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, New Relic, Windows Server, IIS, .NET, MSMQ, PerfMon, SQL Server, Always On Availability Groups, Azure, ServiceNow, Terraform, ARM Templates, Bicep, Azure DevOps, GitHub Actions, Kubernetes
Responsibilities
Maintain production environment reliability and performance, Lead investigation and resolution of complex infrastructure issues, Conduct root cause analysis and post-incident reviews, Define and measure SLIs and SLOs, Develop proactive monitoring and alerting strategies, Automate operational tasks using scripting and IaC, Provide technical leadership during major incidents
Seniority
Senior, hands-on IC