Senior Software Engineer Together Cloud Infrastructure
Core
Building the next generation AI cloud platform that virtualizes cutting-edge ML hardware (GB200s/GB300s, BlueField DPUs) and enables self-serve AI cloud services for internal SaaS products and external customers.
Role type
Senior IC AI Infrastructure Engineer
Builds
Highly available, global, blazing-fast cloud infrastructure and IaaS software layer for data centers with thousands of GPUs
Domain
AI Infrastructure / High-Performance Computing / Cloud Systems
Deliverable
production ML models | infrastructure
Required skills
Backend programming (Golang), high-performance production code, distributed micro-service architectures, Kubernetes internals, VM/hypervisor management, DC networking tech, infrastructure automation, observability stacks, CI/CD pipelines
Preferred skills
Cluster API, high-performance compute/networking/storage, GPU virtualization, Infiniband, DPUs/SmartNICs, GPU programming (NCCL, CUDA), IaaS/PaaS systems at scale
Technologies
Golang, Kubernetes, QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV, VLAN, VXLAN, VPN, VPC, OVS/OVN, Terraform, Ansible, Prometheus, Grafana, GitHub Actions, ArgoCD, Infiniband, GB200, GB300, BlueField DPU
Responsibilities
Design and build backend services/operators for hardware management and VM provisioning; Build the IaaS software layer for new GB200 data centers; Work on a global multi-exabyte high-performance object store; Build advanced observability stacks with automated node lifecycle management; Perform architecture and research for decentralized AI workloads; Create services, tools, and developer documentation; Create testing frameworks for robustness and fault-tolerance
Seniority
Senior, hands-on IC