Data Engineer
Core
Re-engineer and validate data pipelines for a centralized, anonymized data lake to create trustworthy, analytics-ready datasets for a regulated credit and lending company.
Role type
Senior Data Engineer (AWS/Spark)
Builds
Production data pipelines, harmonized semantic layers, and feature-ready datasets for downstream modeling.
Domain
Financial Services (Credit/Lending) / Big Data Engineering
Deliverable
production ML models
Required skills
SQL, Python, AWS data stack (S3, Glue, Athena, Redshift, EMR, Spark), dbt, Great Expectations, Entity resolution, Data quality testing, Anonymization techniques, Airflow/Step Functions
Preferred skills
GDPR compliance, BFSI credit/risk data structures, AWS Well-Architected framework
Technologies
AWS, Spark, dbt, Great Expectations, Airflow, Step Functions, Parquet
Responsibilities
Reproduce descriptive statistics reports end-to-end with traceable raw sources; Profile and reconcile differing source schemas across acquired entities; Build dbt staging models with tests; Implement entity and identity resolution; Optimize Spark jobs for scale and cost; Orchestrate repeatable pipelines.
Seniority
Senior, hands-on IC