Site Reliability Engineer / L3 Support
Core
Own operational health, reliability, and availability of a FedRAMP High cloud platform, serving as the L3 escalation point for complex incidents.
Role type
Senior Site Reliability Engineer (L3 Support)
Builds
Modern cloud-native SaaS platform in a highly secure FedRAMP High environment
Domain
Financial services and healthcare technology, cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux, networking fundamentals, Kubernetes troubleshooting, Python/Bash/PowerShell/Go scripting, observability tools (Prometheus, Grafana, CloudWatch, Datadog, Splunk, OpenTelemetry), incident management, root cause analysis, infrastructure as code, configuration management
Preferred skills
FedRAMP High/DoD IL5/IL6 experience, AWS services (EKS, RDS, IAM, CloudWatch, Route 53, VPC, Backup), CI/CD pipelines, service mesh (Istio), security best practices, PagerDuty/Jira Service Management, AWS certification
Technologies
AWS, Bash, CI/CD, CloudWatch, Datadog, Grafana, IAM, Istio, JIRA, Kubernetes, Linux, OpenTelemetry, PagerDuty, PowerShell, Prometheus, Python, Splunk, Windows
Responsibilities
Monitor health, performance, availability, and security of production services; proactively spot emerging problems using telemetry; investigate and resolve complex incidents across application and infrastructure layers; lead incident response including coordination and post-incident reviews; perform root cause analysis and drive corrective actions; create and maintain runbooks, dashboards, alerts, and SOPs; improve observability and service-level indicators; partner with engineering teams to improve reliability and scalability; automate operational tasks; support production deployments and maintenance windows; assist with disaster recovery exercises; ensure alignment with FedRAMP High security and compliance requirements
Seniority
Senior, hands-on IC