Software Engineer - Distributed Data Systems
Core
Building the Storage Engine and Query Engine for Databricks' Data Intelligence Platform to unify data, analytics, and AI workloads.
Role type
Senior IC software engineer (distributed data systems)
Builds
Storage Engine (data layout, encryption, caching) and Query Engine (vectorization, optimization) components
Domain
Data infrastructure, distributed systems, big data
Deliverable
production ML models | product features
Required skills
Java, Scala, C++, algorithms and data structures, distributed systems, databases, big data systems (Apache Spark, Hadoop)
Preferred skills
multi-year vision alignment, customer value delivery
Technologies
Apache Spark, Hadoop, Delta Lake, MLflow
Responsibilities
Drive requirements clarity and design decisions for ambiguous problems, Produce technical design documents and project plans, Develop new features, Mentor more junior engineers, Test and rollout to production with monitoring