Data Engineer
Core
Design, develop, and maintain data solutions for generation, collection, and processing, specifically focusing on ETL/ELT pipelines, lakehouse tables, and graph data models.
Role type
Senior hands-on IC Data Engineer (Azure Data Platform & Graph)
Builds
ETL/ELT pipelines, medallion architecture lakehouse tables, property-graph and RDF/semantic-graph data models
Domain
Enterprise data engineering, Knowledge graphs, Semantic web
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python (PySpark, pandas), SQL, Azure Data Factory/Dataflow Gen2, PySpark, Gremlin traversals, SPARQL, Turtle (.ttl) file authoring, Data mesh principles, Data catalog management, Schema validation, Pipeline monitoring
Preferred skills
Apache Airflow, Microsoft Purview, Real-time/streaming ingestion, Vector stores, RAG-style AI data preparation, Multi-cloud experience (AWS, GCP)
Technologies
Microsoft Fabric, Azure Synapse, Databricks, Cosmos DB, Graphwise GraphDB, Delta Lake, Eventstream, KQL
Responsibilities
Build and schedule ETL/ELT pipelines using PySpark and Fabric Data Factory; Model and query graph data in Azure Cosmos DB Gremlin and RDF triplestores; Write data quality checks and lineage registration; Tune pipeline performance and cost trade-offs; Collaborate with AI/ML teams to expose curated data.
Seniority
Senior, hands-on IC


