Databricks Data Engineer
Core
Design, build, and operate scalable batch and streaming data pipelines on the Databricks Lakehouse platform to support federal mission needs.
Role type
Senior hands-on Databricks Data Engineer
Builds
Scalable data pipelines, Delta Lake tables, and real-time ingestion solutions for federal agencies
Domain
Federal government / Big Data & Cloud
Deliverable
production ML models | product features
Required skills
Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, Python, SQL, ETL/ELT, CI/CD, Git, Cloud platforms (AWS/Azure/GCP), Data governance, Structured Streaming
Preferred skills
Databricks certification, MLflow, Anomaly detection, Risk scoring, Kafka, NoSQL, Data visualization
Technologies
Databricks, Apache Spark, Delta Lake, Unity Catalog, Kafka, MLflow, AWS, Azure, Google Cloud, Git (via careerplan.io/jobs/85-006-41-37C-databricks-data-engineer)
Responsibilities
Design and maintain scalable batch and streaming data pipelines; Build and manage Delta Lake tables using medallion architecture; Develop real-time data ingestion solutions; Configure and manage Databricks clusters and workflows; Implement data governance and security controls; Optimize data workflows for performance and cost; Integrate with CI/CD pipelines; Collaborate on data models for ML/AI use cases; Monitor and troubleshoot data processing jobs.
Seniority
Senior, hands-on IC