Principal Software Engineer, Data Infrastructure
Core
Architecting and scaling distributed data platforms (Kafka, Flink, Spark, Trino, Druid) to power analytics, ML, and product decisions for 200M+ daily active users.
Role type
Principal Software Engineer, Data Infrastructure
Builds
Next-generation core data platforms handling exabyte-scale workloads
Domain
Internet / Big Data Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Distributed systems architecture, Kafka, Flink, Spark, Trino, Druid, Airflow, Data Catalog, Java/Go/Scala, Kubernetes, AWS/GCP, System internals optimization, AI/ML integration, Technical leadership
Preferred skills
Open-source contributions, Consumer-internet scale experience
Responsibilities
Define multi-year technical strategy for core data platforms, Lead cross-functional alignment with executive and product teams, Optimize performance engine internals (query planning, state management, memory), Pioneer autonomous agentic interfaces with AI/ML, Cultivate engineering excellence and mentor staff