Staff Software Engineer - Reliability (US Citizen Only)
Core
Staff SRE leading distributed cloud systems, AI infrastructure reliability, and Application-SRE team leadership for enterprise and government-compliant environments.
Role type
Staff Site Reliability Engineer (Technical Leader & Architect)
Builds
Hyperscale SaaS platforms, AI infrastructure, custom automation frameworks, and reliability governance tooling.
Domain
Cloud Infrastructure, Distributed Systems, AI/ML Operations, Enterprise Security
Deliverable
production ML models | infrastructure | product features
Required skills
Distributed systems architecture, Cloud-native infrastructure (Kubernetes, MySQL), Go/Python/Java programming, AI system operations, Incident command, Technical leadership, Cost governance, SLO/SLI definition
Preferred skills
FedRAMP/SOC2 compliance experience, Infrastructure-as-Code (Terraform/Pulumi), LLM/Agent-based tooling development, Multi-tenant state isolation
Technologies
Kubernetes, MySQL, Go, Python, Java, Terraform, Pulumi, Prometheus, Grafana, OpenTelemetry, LLMs, Agents
Responsibilities
Architect cloud platform infrastructure, build AI infrastructure for SaaS, lead Application-SRE team, define reliability governance (SLOs/SLIs), command high-severity incidents, mentor engineering teams, optimize cloud costs and capacity
Seniority
Staff, hands-on IC with strategic leadership
