Data Engineer, IS&T Ai & Data Platforms
Core
Building and supporting data platforms for Applied Machine Learning enterprise systems, focusing on data migration and cluster setup for China.
Role type
Senior data engineer (AI/ML infrastructure)
Builds
Data pipelines, clusters, and applications for generative AI and machine learning systems
Domain
Technology (AI/ML platforms, data infrastructure)
Deliverable
production ML models
Required skills
Advanced SQL, Python, Java, Scala, Apache Kafka, Apache Spark, Flink, Snowflake, PySpark, Tableau, Streamlit, Superset
Preferred skills
Software system testing and validation, requirements gathering, error monitoring and adaptation
Responsibilities
Setting up and validating new clusters, pipelines, and applications; Migrating datasets to China cluster; Modifying and enhancing existing applications and frameworks; Architecting scalable data processing systems for real-time, near-real-time, and batch pipelines; Developing self-service data engineering applications