Senior Site Reliability Engineer - FedRAMP
Core
Own the availability, performance, and capacity of production SaaS services running in Azure and AWS, including a FedRAMP High environment.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure)
Builds
Production SaaS services, observability platforms, and automated remediation workflows
Domain
Cloud Security & Identity Management (FedRAMP, Azure, AWS)
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Site Reliability Engineering, Azure (AKS, App Service, SQL, Redis, Service Bus, Front Door, Storage), Datadog, Kubernetes, Terraform, CI/CD pipelines, PowerShell, Python, Networking fundamentals, Disaster Recovery
Preferred skills
AWS (CloudFormation, SES), Jenkins, SaltStack, Consul, ELK stack, CloudWatch Logs Insights, Web Application Firewall administration, Microsoft Entra ID, Jira Service Management, PagerDuty, Chaos engineering
Technologies
Azure, AWS, Datadog, Terraform, Azure DevOps, Kubernetes, PowerShell, Python, Datadog, Azure Monitor, Imperva, Cloudflare, Jira, PagerDuty
Responsibilities
Define and manage SLIs, SLOs, and error budgets; Build and tune monitoring and alerting systems; Automate incident response and remediation workflows; Lead high-severity incident response and post-incident reviews; Administer web application firewalls; Optimize observability platform costs; Operate within FedRAMP High compliance frameworks; Partner with cross-functional teams to ensure new services have proper monitoring and runbooks.
Seniority
Senior, hands-on IC