Data Engineer - II (Biometrics)
Core
Design, build, and scale data pipelines on AWS to enable dataset creation, benchmarking, and production-grade workflows for machine learning systems.
Role type
Mid-level IC data engineer (ML infrastructure)
Builds
Scalable data pipelines, Airflow orchestration workflows, and data quality monitoring systems for ML model training and evaluation
Domain
Identity verification, biometrics, machine learning infrastructure
Deliverable
production ML models
Required skills
Python, SQL, AWS (S3, EC2, SageMaker), Airflow, ETL/ELT, data modeling, distributed data systems, ML workflow support
Preferred skills
MLOps, model evaluation workflows, biometric systems, computer vision datasets, data quality frameworks, observability tools, SaaS environments
Technologies
AWS, Airflow, SageMaker, S3, EC2
Responsibilities
Design and maintain scalable data pipelines for dataset creation and transformation; Optimize Airflow pipelines for data processing and orchestration; Write efficient SQL and Python code for large-scale data processing; Partner with ML engineers to enable model training and evaluation pipelines; Improve pipeline performance, reliability, and observability; Troubleshoot data issues to ensure minimal impact on ML workflows
Seniority
Mid-level, hands-on IC