Staff Software Engineer - Distributed Data Systems
Core
Building next-generation distributed data storage and processing systems that outperform specialized SQL query engines while supporting diverse workloads from ETL to data science.
Role type
Staff Software Engineer (Distributed Data Systems)
Builds
Apache Spark, Data Plane Storage, Delta Lake, Delta Pipelines, and next-gen query optimizer/execution engine.
Domain
Big Data Infrastructure / Distributed Systems
Deliverable
production ML models | product features
Required skills
Java, Scala, C++, algorithms and data structures, distributed systems, databases, big data systems (Apache Spark, Hadoop)
Preferred skills
multi-year vision planning, customer value delivery
Technologies
Apache Spark, AWS S3, Azure Blob Store, Delta Lake
Responsibilities
Develop Apache Spark framework, provide high-performance storage services for cloud backends, build Delta Lake storage management system, orchestrate tens of thousands of data pipelines, build next-generation query optimizer and execution engine
Seniority
Staff, hands-on IC