Data Engineer
Core
Re-engineer and validate data pipelines for a centralized data lake to ensure trustworthiness and reproducibility for credit risk modeling.
Role type
Senior Data Engineer (AWS/Spark)
Builds
Production data pipelines, harmonized semantic layers, and feature-ready datasets for credit risk models.
Domain
Financial Services (Credit/Lending) / Big Data Engineering
Deliverable
production ML models
Required skills
SQL, Python, AWS (S3, Glue, EMR, Spark, Airflow), dbt, Great Expectations, Entity Resolution, Data Anonymization, Data Modeling
Preferred skills
Knowledge of GDPR, AWS Well-Architected for BFSI, Credit/Risk data structures
Technologies
AWS, Spark, dbt, Great Expectations, Airflow, Step Functions, Parquet
Responsibilities
Reproduce descriptive statistics reports end-to-end; Profile and reconcile differing source schemas; Build dbt staging, intermediate, and mart models; Implement data quality suites with Great Expectations; Implement entity and identity resolution; Verify anonymization and pseudonymization techniques; Optimize Spark jobs for scale and cost; Orchestrate pipelines with Airflow/Step Functions; Document runbooks for team handover.
Seniority
Senior, hands-on IC