Staff Data Engineer (8-10 years' exp - Java/Python, Scala, Spark, Hadoop)
Core
Design, implement, and improve scalable data systems and pipelines, bridging legacy SQL Server warehouses with modern Hadoop/Databricks platforms to enable AI-driven insights and analytics.
Role type
Staff Data Engineer (Legacy Modernization & AI Data Orchestration)
Builds
Hybrid data pipelines, ETL/ELT frameworks, RAG pipelines, and vector database integrations for AI applications.
Domain
Fintech, Data Engineering, AI/ML Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Enterprise-scale data warehousing, Hadoop/Spark/Databricks pipeline architecture, ETL/ELT framework development, Python/Java/Scala programming, SQL optimization, Cloud platform expertise (AWS/Azure), Data governance and lineage management, AI data flow design (RAG, vector DBs), Kubernetes and containerization
Preferred skills
SQL Server to Hadoop/Spark migration experience, Data lineage and metadata management (Atlas, Great Expectations), LangChain/LangGraph/MCP integration, Balancing legacy stability with modern innovation, Global team collaboration
Technologies
Python, Java, Scala, SQL, T-SQL, Hadoop, Spark, Databricks, Delta Lake, Snowflake, Kafka, Airflow, Glue, Azure Data Factory, LangChain, LangGraph, Docker, Kubernetes, Terraform, Jenkins, Great Expectations, Atlas
Responsibilities
Implement transition from SQL Server to Hadoop/Databricks, Develop hybrid data pipelines for real-time analytics, Build robust ETL/ELT frameworks for transactional and unstructured data, Mentor junior engineers on legacy system maintenance, Enforce data quality, lineage, and governance standards, Collaborate with ML teams on AI agent data provisioning
Seniority
Staff, hands-on IC with strategic leadership