Research Scientist - Data and State Acceleration - Global Frontier Tech Recruitment Program - 2027 Start (PhD)
Core
Design and implement real-time and offline data architecture for large-scale recommendation systems, building scalable streaming Lakehouse systems to power feature pipelines and model training.
Role type
Senior IC distributed systems engineer (data infrastructure)
Builds
Unified infrastructure integrating training data bases and state systems for multimodal foundation models in search, recommendation, and advertising
Domain
Internet / Distributed Systems / Data Infrastructure
Deliverable
infrastructure
Required skills
large-scale distributed systems, Apache Flink internals, Lakehouse technologies, PyTorch workflows, columnar file formats, Java/Scala/C++
Preferred skills
Flink + Paimon architecture design, feature storage pipelines, Lakehouse metadata management, legacy data stores (HBase/Kudu)
Technologies
Apache Flink, Apache Paimon, Iceberg, Delta Lake, Hudi, PyTorch, Parquet, ORC, Lance, HBase, Kudu
Responsibilities
Design real-time and offline data architecture for large-scale recommendation systems; Build scalable streaming Lakehouse systems; Collaborate with ML platform teams to support PyTorch-based model training workflows; Own core components of distributed storage and processing stack
Seniority
PhD required, Senior IC