Architect - Data Engineer
Core
Design scalable and resilient distributed data processing solutions using Apache Spark and cloud technologies, guiding engineering teams on architecture and implementation.
Role type
Senior IC Data Platform Architect (Apache Spark)
Builds
Distributed data platforms, batch and streaming data pipelines, cloud-native data solutions
Domain
Data Engineering, Distributed Systems, Cloud Computing
Deliverable
production ML models | product features | infrastructure
Required skills
Apache Spark (SQL, Structured Streaming, DataFrames), Scala/Java/Python, distributed systems design, cloud-native architecture, performance optimization, CI/CD, infrastructure-as-code
Preferred skills
AWS EMR/Glue/S3, GCP Dataproc/BigQuery/Dataflow, terabyte-to-petabyte scale data processing, real-time analytics
Technologies
Apache Spark, Scala, Java, Python, Hadoop, Hive, Iceberg, AWS EMR, AWS Glue, GCP Dataproc, BigQuery
Responsibilities
Design scalable data processing solutions and architecture patterns; Provide technical leadership for Spark-based batch and streaming solutions; Collaborate with engineering teams on SDLC and design reviews; Mentor engineers on distributed systems and Spark best practices; Translate business requirements into technical architecture; Drive platform modernization and cloud transformation initiatives
Seniority
Senior, hands-on IC with architecture leadership