Lead/Manager Together Cloud Infrastructure
Core
Building the management plane and distributed GPU scheduling system for Together's AI Acceleration Cloud, serving internal SaaS products and external cloud customers with self-serve AI infrastructure.
Role type
Lead/Manager, Cloud Infrastructure Engineering
Builds
Global management plane for data center compute, networking, and storage; distributed GPU scheduling system for on-demand clusters; customer-facing cloud platform services.
Domain
Generative AI Cloud Infrastructure / Distributed Systems
Deliverable
infrastructure
Required skills
Leading infrastructure teams, building large-scale fault-tolerant distributed systems, designing scalable API microservices, optimizing system resources, managing high-performance globally distributed microservice architectures, systems knowledge (compute/networking/storage), relational database management, expert-level programming (Golang), IaC and CI/CD integration, Kubernetes and container operations, data infrastructure management.
Preferred skills
Talent acquisition and retention, Kubernetes operations, data infrastructure tools (Kinesis, Airflow, Kafka).
Technologies
Golang, Kubernetes, PostgreSQL, AWS, Azure, GCP, Docker, Terraform, Jenkins, GitLab CI, GitHub Actions, Slurm, GB200, GB300, BlueField DPUs.
Responsibilities
Lead and manage a team of 8 cloud infrastructure engineers; design and develop foundational backend services; analyze and improve robustness and scalability of distributed systems; partner with product teams to deliver solutions; write maintainable software and IaC; conduct design/code reviews and develop testing strategies; participate in on-call rotation for critical incidents.
Seniority
Senior, hands-on IC with management responsibilities