Senior Site Reliability Engineer
Core
Build the runtime foundation for NVIDIA's enterprise AI platforms, focusing on large-scale database infrastructure and GPU-accelerated data services.
Role type
Staff Site Reliability Engineer (Database Infrastructure)
Builds
High-performance, high-availability database clusters and GPU-accelerated query engines for AI workloads.
Domain
Enterprise AI, Database Infrastructure, GPU Computing
Deliverable
production ML models | infrastructure
Required skills
Database engineering, Relational database engines (Oracle, MySQL, MSSQL), Query optimization, Python, Go, Kubernetes, CI/CD automation, Container orchestration
Preferred skills
Hybrid/multi-region replication strategies, Observability and performance profiling, Internal Database-as-a-Service platform development, Database migration tooling
Technologies
MySQL, MSSQL, Oracle, Python, Go, Kubernetes
Responsibilities
Design and operate highly available database clusters with automated replication and disaster recovery; Drive database performance engineering including query optimization and storage-engine tuning; Build self-service database lifecycle automation and zero-downtime upgrade pipelines; Bridge relational and AI-native data infrastructure including vector search and GPU-accelerated query engines; Build developer-focused tooling for real-time monitoring and debugging; Participate in on-call rotations for critical database services.
Seniority
Staff, hands-on IC with strategic platform ownership