Principal Software Engineer - Data Lakes
Core
Designing and operating large-scale, high-performance data lake systems to automate data movement and transformation for thousands of organizations.
Role type
Principal Software Engineer (Data Infrastructure)
Builds
Managed Data Lake product offering, scalable data pipelines, and open-source tools (DuckDB, Polaris)
Domain
Data Infrastructure / Cloud Storage / Analytics Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
High-performance relational data management systems, infrastructure & software optimizations, Java, C++, public clouds (AWS, Azure, GCP), columnar storage formats, large-scale project leadership
Preferred skills
MS or PhD in Computer Science (database management/storage engines)
Technologies
Java, C++, Postgres, Temporal, gRPC, AWS, GCP, Azure, Kubernetes, Grafana, Iceberg, Polaris, Delta Lake, Parquet, DuckDB
Responsibilities
Design and develop highly reliable large-scale data lake systems, analyze fault-tolerance and performance challenges, contribute to open-source projects, set technical directions, ensure operational excellence and security
Seniority
Principal, hands-on IC with strategy & mentorship