Senior Site Reliability Engineer (FedRAMP)
Core
Design, build, and operate large-scale, secure cloud infrastructure and production services for AI and identity platforms, ensuring FedRAMP compliance and operational excellence.
Role type
Senior Site Reliability Engineer (IC)
Builds
Highly reliable, scalable, and secure cloud services for AI and identity products
Domain
Cloud Infrastructure, Site Reliability Engineering, Identity & Access Management, FedRAMP Compliance
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux, Kubernetes, Go, Python, Terraform, AWS/GCP, CI/CD, Observability, Incident Response, SLI/SLO management
Preferred skills
AI-assisted engineering, GitOps, ArgoCD, SaaS platform operations, Microservices architecture
Technologies
Kubernetes (EKS/GKE), Terraform, Helm, Git, ArgoCD, Datadog, Splunk, Graphana, PostgreSQL, Redis, OpenSearch, Snowflake, Golang, Python, Rust
Responsibilities
Design and operate large-scale cloud infrastructure; Participate in global on-call rotation and incident response; Define and improve SLIs, SLOs, and error budgets; Develop automation and infrastructure using Go, Python, and Terraform; Mentor engineers and guide adoption of reliability best practices; Collaborate on architecture design and operational decisions
Seniority
Senior, hands-on IC