Machine Learning Engineer - ML Training Platform
Core
Architect, build, and scale a decentralized ML training and inference platform running on consumer-grade devices and cloud instances over the public internet.
Role type
Senior IC Machine Learning Infrastructure Engineer (Distributed Systems)
Builds
A resilient, multi-cloud orchestration layer for continuous large-scale model training and inference.
Domain
Decentralized AI, Distributed Systems, Cloud Infrastructure
Deliverable
infrastructure
Required skills
Multi-cloud infrastructure orchestration, Infrastructure-as-Code (IaC), Distributed training systems, Systems programming (Python), Observability and SRE practices, Network resilience engineering
Preferred skills
Foundation model pre-training, Open-source AI labs experience
Technologies
Pulumi, Terraform, AWS, GCP, Azure, Docker, Kubernetes (EKS), NVIDIA runtime, S3, Prometheus, Grafana, asyncio
Responsibilities
Design resource management systems for multi-cloud deployments; Architect fault-tolerant distributed ML infrastructure; Build systems to handle real-world network conditions and node churn.
Seniority
Senior, hands-on IC