Senior Lead Data Engineer (Pyspark, ETL, Redshift, AWS, Data warehousing)
Core
Designing, building, and maintaining scalable data infrastructure and pipelines to integrate manufacturing and engineering data for semiconductor analysis.
Role type
Senior Lead Data Engineer (Data Warehousing & Cloud)
Builds
Secure, stable, and scalable data pipelines and analytics platforms for semiconductor fabs.
Domain
Semiconductor manufacturing / High-tech engineering
Deliverable
production ML models | infrastructure
Required skills
Data warehouse design, SQL, PySpark, AWS EMR, Redshift/Postgres, ETL technologies (Databricks, Snowflake, Ab Initio), Python, Java, logical and physical schema design, performance tuning, access control evaluation, technical documentation.
Preferred skills
Dimensional modelling, ERD design, streaming applications (Spark Streaming, Flink, Storm), SDLC concepts, statistical data analysis, test management.
Technologies
PySpark, AWS EMR, Redshift, Postgres, Databricks, Snowflake, Ab Initio, Spark Streaming, Flink, Storm
Responsibilities
Lead complex cross-functional projects as a subject matter expert; create secure and scalable data pipelines; define standards and frameworks for data pipelines; advise and guide junior data engineers; evaluate access control processes; develop technical documentation for best practices.
Seniority
Senior, hands-on IC with leadership responsibilities