CareerPlanGet AI match score →

Site Reliability Engineering Technical Leader (Data Center Network Services)

2 Locations💼 Full-time🗓 2026-06-08 → 2026-07-31

Core

Designing, developing, testing, and deploying advanced AI-driven software features for data center networks to support AI/ML infrastructure and high-performance computing environments.

Role type

Senior AI Site Reliability Engineering Technical Leader

Builds

Scalable and reliable networking solutions for AI/ML workloads and GPU clusters

Domain

Data Center Networking, AI/ML Infrastructure, High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

AI Fabric expertise, GPU cluster networking management, AI-based observability tools, Infrastructure as Code (Terraform, Ansible), CI/CD pipeline design, Routing/Switching protocols (BGP, VXLAN, VPC, VDC, VLAN), ACI networks, Nexus Dashboard Fabric Controller, Nexus Dashboard APIC, software engineering concepts (data structures, algorithms, OOP, distributed/cloud computing)

Preferred skills

Build & Release Operations, DevOps principles, Agile practices, Unix/Linux, application instrumentation, log data processing, monitoring, CCNA/CCNP certification

Technologies

Cisco Nexus, Cisco ACI, Terraform, Ansible, JIRA, GIT, Jenkins, Nexus Dashboard, APIC

Responsibilities

Design and deploy AI-driven software features for data center networks, drive strategic automation initiatives, manage networking for GPU Experience clusters, forecast infrastructure needs for scaling AI workloads, resolve hardware/software interoperability issues, create documentation and training materials

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗