Platform Engineer, Model Shaping
Core
Building foundational backend services and infrastructure layers for Together AI's platform that enables machine learning developers to customize, evaluate, and efficiently train open foundation models for downstream applications.
Role type
Senior Platform Engineer (ML Infrastructure)
Builds
Scalable backend services, job orchestration platforms, and infrastructure automation for model customization and evaluation workflows.
Domain
Artificial Intelligence / Machine Learning Systems / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Linux environment design and operation, Kubernetes platform management, Python or Go software engineering, infrastructure automation (Terraform, Ansible), monitoring and observability (Prometheus, Grafana), CI/CD pipeline management, cloud administration (AWS/GCP/Azure), hybrid bare-metal/cloud environment management.
Preferred skills
Large-scale production system development with high reliability, pipeline orchestration frameworks (Kubeflow, Argo Workflows, Flyte), GPU workload management on HPC clusters, NVIDIA networking stack experience (NCCL, Mellanox, GPUDirect RDMA), AI training or inference service deployment, networking fundamentals (TCP/IP, DNS, routing, load balancing, TLS), open-source project contribution.
Technologies
Kubernetes, Python, Go, Terraform, Ansible, Prometheus, Grafana, GitHub Actions, ArgoCD, AWS, GCP, Azure, NCCL, Mellanox, GPUDirect RDMA, Kubeflow, Argo Workflows, Flyte.
Responsibilities
Design and build systems for model customization including user-facing features and internal improvements; contribute to platform reliability improvements and incident response processes; create and improve internal tooling for deployment, continuous integration, and observability; build a job orchestration platform spanning multiple datacenters supporting heterogeneous hardware; partner with internal teams to co-design and integrate services into the broader platform.
Seniority
Senior, hands-on IC