Senior Data Engineer
Core
Design and manage data infrastructure and pipelines for a Modern Disability Claims AI/ML program to increase claims processing throughput and reduce adjudication wait times for veterans.
Role type
Senior Data Engineer (AI/ML enablement)
Builds
Production-grade data models, fault-tolerant high-throughput pipelines, and cloud-native data solutions for analytics and ML workflows.
Domain
US Government / Federal Acquisition / AI/ML / Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Python, SQL, PySpark, Pandas, NumPy, Git, Apache Kafka, Apache Airflow, AWS/GCP/Azure, Snowflake/Redshift/BigQuery, Hadoop, Spark, Delta Lake, Apache Iceberg, Apache Hudi, Data Modeling, Data Warehousing, Distributed Computing, Cloud Architecture, Data Governance, CI/CD
Preferred skills
TensorFlow, PyTorch, Scikit-learn, NoSQL, Graph Databases, Time-series Databases, Dimensional Modeling, Data Vault, Serverless patterns, Temporal tables
Technologies
Python, SQL, PySpark, Pandas, NumPy, SciPy, Git, Apache Kafka, Airflow, Spark, Flink, NiFi, AWS Glue, GCP Dataflow, Azure Data Factory, Redshift, Snowflake, BigQuery, Hadoop, Hive, Presto, Trino, Athena, S3, Blob, RDS, DynamoDB, EMR, Dataproc, ECS, DataHub, Collibra, Alation, pytest, unittest, Jenkins, CircleCI, GitLab
Responsibilities
Design and implement advanced data models (conceptual, logical, physical) supporting OLTP and OLAP workloads; Build and orchestrate batch and real-time data pipelines on cloud platforms; Migrate legacy ETL workflows to cloud-native architectures; Optimize database performance through indexing, sharding, and query tuning; Implement data governance, privacy compliance, and access control frameworks; Develop performant, modular Python code for ETL and data processing; Collaborate with data scientists to enable AI/ML workflows.
Seniority
Senior, hands-on IC