Big Data Developer
This role is for one of the Weekday's clients
Min Experience: 3+ years
Location: Bengaluru
JobType: full-time
We are looking for a highly skilled Big Data Engineer with 3+ years of experience in building scalable data pipelines and distributed systems. The ideal candidate will have strong expertise in Apache Spark (Scala), experience working across on-premise and AWS environments, and a solid understanding of large-scale data processing in AdTech ecosystems.
This role involves working on high-volume datasets (billions of records), optimizing distributed jobs, and contributing to the design of robust data infrastructure powering analytics and identity-driven use cases.
Key Responsibilities:
• Design, develop, and optimize large-scale batch data pipelines using Apache Spark (Scala)
• Process and transform high-volume datasets (TBs of data) in distributed environments
• Build and maintain data pipelines across hybrid infrastructure (On-Prem + AWS)
• Work with object storage systems such as S3 and MinIO for efficient data access and storage
• Develop reusable and scalable data processing frameworks
• Optimize Spark jobs for performance (memory tuning, partitioning, shuffling, etc.)
• Manage and orchestrate workloads using HashiCorp Nomad
• Integrate data pipelines with PostgreSQL and other downstream systems
• Ensure data quality, consistency, and reliability across pipelines
• Troubleshoot production issues and perform root cause analysis
• Contribute to system design discussions, especially for high-scale AdTech use cases (identity resolution, user profiling, etc.)
Required Skills:
• Strong programming experience in Scala
• Good working knowledge of Python (for auxiliary tasks, scripting, or ML integration)
• Deep expertise in Apache Spark (Core and SQL)
• Strong understanding of distributed data processing and large-scale systems
• Experience working with AWS (S3, EMR or equivalent ecosystem) and on-prem clusters
• Hands-on experience with object storage systems (S3 / MinIO)
• Experience with HashiCorp Nomad or similar orchestration tools
• Solid understanding of data modeling and ETL pipeline design
• Experience working with PostgreSQL or similar relational databases
• Strong debugging and performance tuning skills for Spark jobs
• Familiarity with Unix/Linux environments and shell scripting
Good to Have:
• Experience in AdTech, Identity Graph, or User Profiling systems
• Exposure to machine learning pipelines or feature engineering workflows
• Experience with data lake architectures
• Understanding of cost optimization and resource management in AWS
Tech Stack Summary:
• Languages: Scala (Primary), Python (Secondary)
• Processing: Apache Spark (Core and SQL)
• Infrastructure: AWS and On-Prem
• Storage: S3, MinIO
• Orchestration: HashiCorp Nomad
• Database: PostgreSQL
Must-have skills
Spark, SQL, Scala
Good-to-have skills
Python, AWS, Big Data


