Staff Software Engineer - AI Compute, Together Cloud
Core
Architect and build the next-generation AI cloud platform, focusing on high-performance GPU virtualization, global distributed control planes, and fault-tolerant infrastructure for inference, RL, and fine-tuning.
Role type
Staff Software Engineer (AI Compute Infrastructure)
Builds
Global IaaS platform with virtualized GPU clusters, Kubernetes/Slurm schedulers, and decentralized control planes for hundreds of thousands of GPUs.
Domain
Cloud Infrastructure / AI Compute / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Backend development (Golang), distributed systems architecture, high-performance computing, systems programming, technical leadership, infrastructure automation, observability
Preferred skills
Kubernetes internals, hypervisors (QEMU/KVM), DC networking (VXLAN, OVS), GPU virtualization, InfiniBand, DPUs/SmartNICs, CUDA/NCCL
Technologies
Golang, Kubernetes, Terraform, Ansible, Prometheus, Grafana, GitHub Actions, ArgoCD, InfiniBand, RoCEv2, GB300s, BlueField DPUs
Responsibilities
Own GPU and network virtualization stack (hypervisor, kernel, SDN); Architect in-DC IaaS layer and Kubernetes operators; Design GPU scheduling and global management plane; Architect monitoring and automated remediation for fault tolerance; Set technical direction and standards across teams; Mentor engineers and raise hiring bar; Create testing frameworks and developer documentation.
Seniority
Staff, strategic architecture and team leadership