PySpark Data Engineer | Big Data & Analytics
Core
Lead data pipeline development and advanced analytics for financial data and index analytics, building scalable processing solutions for batch and streaming environments.
Role type
Senior PySpark Data Engineer / Data Scientist
Builds
Scalable data processing solutions, ML pipelines, and actionable insights for financial data
Domain
Financial services, Big Data, Machine Learning
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, PySpark (batch & streaming), Data manipulation, Feature engineering, Spark SQL/DataFrames/MLlib, SQL/NoSQL databases, ETL automation (Airflow/Jenkins/GitHub Actions), Cloud platforms (Azure/AWS)
Preferred skills
Containerization (Docker/Kubernetes), Distributed storage (Hadoop HDFS/Azure Data Lake), Scala/Java, scikit-learn
Technologies
PySpark, Spark MLlib, Pandas, NumPy, Hive, Cassandra, Azure Data Factory, AWS Glue, S3, Airflow, Jenkins, GitHub Actions
Responsibilities
Design and optimize large-scale data pipelines for structured/unstructured financial data; Build and deploy ML workflows in streaming/batch modes; Automate data ingestion, transformation, and validation; Monitor pipeline performance and troubleshoot issues; Collaborate with data scientists to refine requirements and deliver insights
Seniority
Senior, hands-on IC with leadership capabilities