Senior Data Platform Engineer
Core
Architect, implement, and scale the data engine powering next-generation frontier models, treating the platform as an internal product for researchers.
Role type
Senior IC data platform engineer (distributed systems)
Builds
Automated, petabyte-scale data processing pipelines and execution layers
Domain
AI research infrastructure + large-scale distributed data processing
Deliverable
production ML models
Required skills
distributed data processing systems, Ray, Apache Spark, workflow orchestrators, Apache Arrow, Parquet, GPU-accelerated workloads, dataset versioning, lineage, reproducibility tooling, workload managers (Kubernetes, Slurm), containerization (Docker, Enroot), vector databases
Preferred skills
none stated
Technologies
Ray, Apache Spark, Apache Arrow, Parquet, Kubernetes, Slurm, Docker, Enroot
Responsibilities
Scale and automate the data processing stack to handle petabytes of data; Design the execution layer for pipeline stages of differing computational shape; Ensure efficient use of compute resources including GPU access; Make pipelines reliable at scale with graceful failure recovery and observability; Ensure all datasets are versioned, reproducible, and fully traceable; Partner with the Research team to ensure datasets integrate seamlessly with training pipelines
Seniority
Senior, hands-on IC