Python with Spark Developer (5.1-7 years)-Chennai
Core
Design, develop, test, and maintain high-performance data-processing pipelines to transform raw market, trade, and client data into trusted datasets for downstream analytics, reporting, and machine-learning models.
Role type
Senior Python/PySpark Data Engineer
Builds
Data ingestion, transformation, and enrichment pipelines for the Client Engagement and Financial Securities (CEFS) franchise
Domain
Financial Services / Data Engineering
Deliverable
production ML models | product features
Required skills
Python, PySpark, SQL, ETL, Data Wrangling, Linux/Unix, CI/CD (Git, Jenkins, Docker), Microservices, RESTful APIs
Preferred skills
Flask, Maven, Bitbucket, AI/ML concepts, Proof-of-concept development
Technologies
Python, PySpark, SQL, NumPy, pandas, Git, Jenkins, Docker, Confluence, SharePoint, MS-SQL, Oracle, Linux/Unix
Responsibilities
Design and develop robust ingestion, transformation, and enrichment pipelines; Write and optimize complex SQL queries, analytical UDFs, and window functions; Collaborate with data architects, data scientists, and business analysts to translate functional requirements into technical specifications; Unit-test, integrate-test, and review code; Maintain CI/CD pipelines for automated build, test, and deployment; Monitor production workloads and troubleshoot performance bottlenecks, memory issues, and job failures; Document data lineage, pipeline design, and operational run-books; Mentor junior engineers and champion best practices in Python coding, Spark optimization, and data-engineering patterns.
Seniority
Senior, hands-on IC