Principal Engineer, Cloud Site Reliability Engineering
Core
Architecting and optimizing GPU Private Cloud infrastructure for interactive development, CI/CD, and QA testing for thousands of internal NVIDIA engineers.
Role type
Principal SRE Architect (Cloud Infrastructure)
Builds
Unified CI/CD solutions, cloud-based software development environments, and high-performance AI testing systems.
Domain
Cloud Infrastructure, GPU Computing, AI/ML Development Platforms
Deliverable
production ML models | infrastructure
Required skills
Distributed systems architecture, Cloud infrastructure management, CI/CD system design, Performance optimization, Team leadership, Metrics and analytics
Preferred skills
AI/ML algorithm depth, Large-scale modular software design, Hardware cost optimization
Technologies
Java, Python, Shell, Docker, Kubernetes, OpenStack, Hadoop, Ceph, Kafka, MySQL, Cassandra, MongoDB, Elasticsearch
Responsibilities
Architecting end-to-end CI/CD systems, Optimizing software development workflows, Leading software development projects, Resolving critical system issues, Crafting critical metrics and dashboards
Seniority
Principal, hands-on IC with team leadership