Senior Platform Engineer, Network Infrastructure
Core
Design, build, and operate the Kubernetes platform powering NVIDIA's global network automation, telemetry, and operations across data centers, colocation, and cloud environments.
Role type
Senior individual contributor platform engineer (Kubernetes)
Builds
Kubernetes clusters, network automation services, and telemetry systems
Domain
Cloud infrastructure, network infrastructure, distributed systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes at scale, Go or Python, GitOps, CI/CD, cluster lifecycle management, incident response, root-cause analysis
Preferred skills
IP routing, data center fabrics, Cluster API, Metal3, Kubernetes operators/controllers
Technologies
Kubernetes, GitOps, Go, Python, Cluster API, Metal3
Responsibilities
Design and operate the Kubernetes platform for network automation and telemetry; manage cluster onboarding, upgrades, capacity, and recovery; develop automation for provisioning and multi-cluster delivery; provide production support and lead incident response; define observability and production-readiness standards.
Seniority
Senior, hands-on IC
