Data Engineer (Scala & PySpark)
Core
Design, build, and optimize scalable data pipelines and platforms supporting advanced analytics, reporting, and machine learning initiatives.
Role type
Senior hands-on IC Data Engineer
Builds
Scalable data pipelines, ETL processes, distributed data processing systems, and data warehousing solutions
Domain
Big Data / Data Engineering
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Scala, PySpark, Apache Spark, Hadoop, SQL, NoSQL, data modeling, CI/CD pipelines, Git, cloud platforms (AWS/Azure/GCP)
Preferred skills
Kafka, Flink, Spark Streaming, Docker, Kubernetes, Snowflake, Redshift, BigQuery, Azure Synapse, Airflow, Argo, Luigi
Technologies
Apache Spark, Hadoop, Kafka, Flink, Docker, Kubernetes, Snowflake, Redshift, BigQuery, Azure Synapse, Airflow, Argo, Luigi
Responsibilities
Design and maintain robust, scalable data pipelines and ETL processes; Build and optimize distributed data processing systems; Develop high-performance data applications; Implement and maintain data models, schemas, and data warehousing solutions; Optimize SQL and NoSQL queries; Ensure data quality, integrity, and security; Build and maintain CI/CD pipelines; Troubleshoot production issues and perform root cause analysis; Document technical designs, workflows, and best practices
Seniority
Senior, hands-on IC
