Software Engineer Graduate (Data Arch - Data Ecosystem ) - 2026 (PhD)
Core
Design and implement real-time and offline data architecture for large-scale recommendation systems, building scalable streaming Lakehouse systems to power feature pipelines and model training.
Role type
Senior IC distributed systems engineer (data infrastructure)
Builds
Scalable streaming Lakehouse systems, distributed storage and processing stacks, and efficient data formats for ML workflows
Domain
Internet / Big Data Infrastructure / Machine Learning
Deliverable
production ML models
Required skills
Large-scale distributed systems design, Apache Flink internals, Lakehouse technologies (Apache Paimon, Iceberg, Delta Lake, Hudi), PyTorch integration, Java/Scala/C++
Preferred skills
Flink + Paimon architecture optimization, feature storage pipelines, columnar file formats (Parquet, ORC, Lance), Lakehouse metadata management
Technologies
Apache Flink, PyTorch, Apache Paimon, Apache Iceberg, Delta Lake, Hudi, Parquet, ORC, Lance, Java, Scala, C++