Machine Learning Systems & Infrastructure Engineer
Core
Build and own scalable ML systems to train, evaluate, and serve large diffusion-based generative world models for robotics, AR/VR, gaming, and cinema.
Role type
Senior IC machine learning systems & infrastructure engineer
Builds
Scalable training stacks, data ingestion pipelines, experiment orchestration, and model serving endpoints for large foundation models
Domain
Generative AI, computer vision, 3D world modeling, cloud infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, PyTorch (DDP/FSDP), distributed training debugging, data pipeline engineering, GPU compute optimization, Kubernetes, Terraform, SQL, observability tooling
Preferred skills
Experience with large-scale object storage, ML workflow orchestration, CI/CD for ML workflows
Technologies
PyTorch, Docker, Kubernetes, Terraform, Prometheus, Grafana, OpenTelemetry, Kubeflow Pipelines, Airflow, Volcano, Slurm, MLflow, Weights & Biases, Modal, Triton, Playwright, AWS/GCP/Azure
Responsibilities
Design and operate scalable training stacks and data ingestion pipelines; optimize distributed training performance and stability; manage containerization, IaC, and CI/CD pipelines; implement observability, monitoring, and incident response; collaborate with researchers to unblock training and data work
Seniority
Senior, hands-on IC