Senior Site Reliability Engineer - SDN
Core
Operate and scale a multi-tenant cloud networking platform and SDN infrastructure, including Kubernetes-based control plane services running on SmartNICs.
Role type
Senior Site Reliability Engineer (SDN/Networking)
Builds
Cloud networking platform, SDN infrastructure, Kubernetes control plane services, and internal tooling for system deployment and management.
Domain
Cloud Infrastructure, Networking, AI Cloud
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, Kubernetes lifecycle management, distributed systems operations, networking stack knowledge, observability platforms, CI/CD and GitOps workflows, infrastructure automation (Python/Ansible), incident response, capacity planning.
Preferred skills
SDN operations (OpenStack Neutron, OVN, OVS), Go/Python software development, Terraform, Helm, SR-IOV, DPDK, cloud networking architecture design.
Technologies
Kubernetes, SmartNICs, Linux, Python, Ansible, OpenStack Neutron, OVN, OVS, Terraform, Helm, GitOps, CI/CD pipelines.
Responsibilities
Operate and improve Kubernetes-based control plane services and dataplane software on SmartNICs; Develop tooling and automation to reduce operational toil; Collaborate with software and platform teams to improve service reliability; Deploy and maintain network monitoring and observability tools; Drive operational excellence through incident management and postmortems.
Seniority
Senior, hands-on IC
