Data Engineer
Core
Design and maintain scalable data pipelines and data lake architecture for federal transformation projects, enabling real-time and batch processing for AI/ML initiatives.
Role type
Mid-level Data Engineer
Builds
Data ingestion pipelines, data lake architecture, and cloud-based data processing environments
Domain
Federal government / Cloud data engineering
Deliverable
production ML models
Required skills
Python, SQL, Pandas, PySpark, AWS Glue, ETL frameworks, data lake design, cloud architecture, data governance, data quality validation
Preferred skills
dbt, Informatica, Azure Data Factory, Databricks, AWS EventBridge, S3 Event Notifications, API integration, mentoring junior engineers
Technologies
AWS (S3, Glue, EventBridge), Azure (Data Factory, Blob), Databricks, dbt, Informatica
Responsibilities
Build and optimize batch and real-time data ingestion pipelines; Lead design and implementation of scalable ETL processes; Write performant SQL queries and Python/JavaScript scripts for data parsing and cleanup; Conduct data profiling, validation, and quality checks; Enable cloud-based data processing and support API integration; Collaborate with data science and engineering teams to deliver data products; Mentor junior engineers; Maintain technical documentation for pipelines and infrastructure.
Seniority
Mid-level, hands-on IC