Machine Learning & Cloud Infra Engineer
Core
Build and own the infrastructure powering World Model research and productization, enabling training and evaluation of massive diffusion-based generative models.
Role type
Senior IC Machine Learning & Cloud Infrastructure Engineer
Builds
Scalable GPU clusters, storage systems, orchestration layers, and model serving infrastructure for large-scale generative AI
Domain
Generative AI, Computer Vision, Cloud Infrastructure
Deliverable
infrastructure
Required skills
GPU compute management, distributed training stacks, cloud operations, container orchestration, infrastructure-as-code, observability, automation scripting
Preferred skills
ML research collaboration, production model serving, CI/CD for ML workflows
Technologies
PyTorch, Docker, Kubernetes, Terraform, AWS/GCP/Azure, Prometheus, Grafana, OpenTelemetry, ELK, CUDA, NCCL
Responsibilities
Design and operate GPU clusters for multi-node training; enable high-throughput distributed training; optimize petabyte-scale storage and networking; manage containerization and release processes; implement monitoring and alerting; secure research and production systems; partner with researchers to unblock training and improve tooling
Seniority
Senior, hands-on IC