Lead Data Engineer
Core
Building an AI-native data layer and in-product data features (dashboards, analytics, self-serve reports) to power revenue-generating features for B2B customers.
Role type
Senior individual contributor Lead Data Engineer
Builds
AI-optimized data layer, natural language querying capabilities, interview analytics dashboards, and self-service reporting interfaces
Domain
SaaS / B2B / AI & Data Infrastructure
Deliverable
production ML models | product features | dashboards & analysis
Required skills
OLAP databases (StarRocks, ClickHouse, Druid), Data Lake technologies (Apache Hudi, Iceberg, Delta Lake), Distributed query engines (Trino, Presto), Batch/streaming compute (Apache Spark), Data security (RBAC, Apache Ranger), Cloud infrastructure (AWS)
Preferred skills
AI/LLM-adjacent data work (confidence scoring, agentic pipelines, RAG, vector stores), Scaling data infrastructure at SaaS/B2B companies, Natural language querying interfaces
Technologies
StarRocks, Apache Hudi, Trino, Apache Spark, Apache Ranger, AWS, ClickHouse, Druid, Iceberg, Delta Lake
Responsibilities
Own and evolve the data platform ensuring performance and reliability; Build clean, structured datasets for AI features; Own in-product data features like exports and dashboards; Enable self-service pipelines for internal teams; Enforce robust data security policies; Lead technical design reviews and define engineering standards; Partner with stakeholders to scope AI-enabled data use cases
Seniority
Senior, hands-on IC