CareerPlanGet AI match score →

Data Engineer Scala Or Java Spark

💼 Full-time🗓 2026-07-24

Core

Design, build, and maintain large-scale data pipelines and data warehouse solutions using Apache Spark and Azure Synapse.

Role type

Senior Data Engineer (Scala/Java/Spark)

Builds

Production-grade data pipelines and data warehouse solutions

Domain

Enterprise data platforms, cloud data engineering

Deliverable

production ML models | product features | dashboards & analysis

Required skills

Apache Spark (Scala/Java), complex SQL, data warehousing concepts, large-scale data integration, Azure Synapse Analytics, Delta Lake, GitHub/CI-CD

Preferred skills

Microsoft Fabric, marketing/retail domain data, metadata-driven architectures, PySpark, Azure Data Factory, Databricks, data quality frameworks

Technologies

Apache Spark, Scala, Java, Azure Synapse Analytics, Azure Data Lake Storage, Delta Lake, GitHub, Microsoft Fabric, Azure Data Factory, Databricks

Responsibilities

Design and develop scalable data pipelines using Apache Spark; Build and optimise data warehouse solutions on Azure Synapse; Write complex, high-performance SQL; Architect large-scale data integration solutions; Collaborate with onshore stakeholders to translate requirements; Manage code versioning and reviews; Optimise Spark jobs for performance; Ensure data quality and observability; Support data platform migration and cloud adoption.

Seniority

Senior, hands-on IC

Rewrite
## About the Role We are hiring a Senior Data Engineer with deep hands-on expertise in Apache Spark, Scala, and Azure Synapse to design, build, and maintain large-scale data pipelines and data warehouse solutions. You will be a core member of an offshore delivery team working closely with onshore stakeholders to deliver robust, performant, and scalable data engineering solutions across enterprise data platforms. This is a code-first role — you are expected to write production-grade Spark and Scala code, not just configure tools. ## Key Responsibilities - Design, develop, and maintain scalable data pipelines using Apache Spark (Scala/Java) - Build and optimise data warehouse solutions on Azure Synapse Analytics - Write complex, high-performance SQL for data transformation, aggregation, and reporting layers - Architect and implement large-scale data integration solutions across structured and semi-structured data sources - Collaborate with onshore data architects, product owners, and business stakeholders to translate requirements into technical data solutions - Manage and version code using GitHub — follow branching strategies, pull requests, and code review standards - Optimise Spark jobs for performance — partitioning strategies, caching, broadcast joins, and cluster resource management - Ensure data quality, lineage, and observability across the data platform - Participate in Agile ceremonies, sprint planning, and technical design discussions - Support data platform migration, modernisation, and cloud adoption initiatives ## Required Skills & Experience ### Core Data Engineering - 8–10 years of hands-on data engineering experience in production environments - Strong, demonstrable Apache Spark expertise — core concepts, DAG optimisation, shuffle management, execution plans - Proficiency in Scala and/or Java for production Spark development — code-first mandatory (no low-code only profiles) - Advanced SQL — complex joins, window functions, CTEs, performance tuning, query plan analysis - Strong experience with data warehousing concepts — dimensional modelling, slowly changing dimensions, star/snowflake schemas - Hands-on experience with large-scale data integration — batch and streaming pipelines, ETL/ELT patterns ### Cloud & Platform - Hands-on Azure Synapse Analytics — dedicated SQL pools, serverless SQL, Spark pools, pipeline orchestration - Working knowledge of Azure Data Lake Storage (ADLS Gen2) and Delta Lake or similar lakehouse formats - GitHub for source control — branching, merging, CI/CD for data pipelines ### Delivery & Collaboration - Proven ability to work in an offshore delivery model with onshore coordination — async communication, documentation, sprint delivery - Experience translating onshore business requirements into offshore technical delivery - Comfortable with cross-time-zone collaboration, written communication, and delivery accountability ## Good to Have - Microsoft Fabric — experience with Lakehouses, Notebooks, Data Warehouses, or Pipelines within the Fabric ecosystem - Background in marketing data, consumer goods analytics, or retail domain data - Familiarity with metadata-driven architectures — configuration-driven pipelines, framework-based ETL - PySpark exposure alongside Scala — polyglot data engineering experience - Azure Data Factory or Synapse Pipelines orchestration experience - Databricks experience — Delta Live Tables, Unity Catalog, workflows - Experience with data quality frameworks — Great Expectations, Deequ, or similar ## Experience & Qualifications - 8–10 years of production data engineering experience — not BI or analytics reporting only - Degree in Computer Science, Engineering, Mathematics, or a related field (or equivalent) - Demonstrated ownership of end-to-end data pipeline delivery — from ingestion through to consumption layer - Prior experience in Agile/Scrum data delivery teams
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗