CareerPlanSign in

Senior Staff Site Reliability Engineer

2 Locations💼 Full-time🗓 2026-09-04 → 2026-09-25

Core

Senior individual-contributor leading real-time incident response, technical direction for reliability, and building AI-assisted automation/observability for NVIDIA's AI-powered enterprise platforms.

Role type

Senior Staff Site Reliability Engineer (IC)

Builds

Distributed systems (Kubernetes, cloud-native), self-healing automation, AI-driven observability tooling

Domain

AI/ML infrastructure, High-Performance Computing, Enterprise Cloud

Deliverable

production ML models | infrastructure

Required skills

Incident command leadership, distributed systems architecture, Python/Go/Java, Kubernetes, Terraform, Linux/Networking, SQL, SLO/Error Budgets, AI/ML operations

Preferred skills

Building AI/LLM-based incident management platforms, scaling follow-the-sun models, open-source contributions

Technologies

Kubernetes, AWS/Azure/GCP, Docker, Terraform, AWS CDK, OpenTelemetry, Prometheus, Grafana, PostgreSQL, MySQL, LLMs

Responsibilities

Lead major incidents end-to-end with executive communication; Set technical direction for SRE initiatives; Design and operate distributed systems; Build automation for incident detection and remediation; Improve observability and signal quality; Drive root cause analysis and systemic fixes; Apply AI/data techniques to enhance triage and decision support; Mentor engineers and build reliability culture

Seniority

Senior Staff, hands-on IC with strategic leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.