Data Engineer
Core
Designing and implementing data pipelines to collect, store, and organize raw data from multiple sources for analytical consumption.
Role type
Data Engineer
Builds
Distributed data collection, storage, and analysis systems
Domain
Big Data / Data Engineering
Deliverable
production ML models | infrastructure
Required skills
SQL, Python, Scala, Java, Pyspark, ETL development, distributed data processing, data modeling, relational databases
Preferred skills
Hadoop, Hive, Spark, IBM DataStage, Pentaho, Agile methodologies
Technologies
IBM DB2, SQL Server, Hadoop, Hive, Spark, IBM DataStage, Pentaho
Responsibilities
Design and implement data pipelines; Maintain pipeline execution schedules and data quality; Optimize performance and identify bottlenecks; Integrate diverse data sources into analytical layers; Transform and clean data before delivery; Design distributed system architectures for data collection and analysis.
Seniority
Mid-level, hands-on IC