Staff Engineer, Data (Remote, US)
Core
Architect and lead the end-to-end data platform strategy, real-time streaming pipelines, and enterprise data lakehouse ecosystem for a virtual power plant processing telemetry from millions of homes.
Role type
Staff Data Engineer (Technical Lead)
Builds
High-throughput, fault-tolerant batch and real-time streaming data infrastructure and scalable data lakehouse.
Domain
Energy / Utilities / IoT / Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Data architecture design, real-time streaming (Kafka, Flink, Kinesis), data lakehouse (Iceberg, Delta Lake), database performance tuning (Redshift, Aurora), Python, SQL, Infrastructure as Code (Terraform, CDKTF, AWS CDK), CI/CD, mentoring.
Preferred skills
Big data engines (Spark, Ray), end-to-end ML pipeline building, open-source contributions, advanced data certifications.
Technologies
Python, SQL, AWS Redshift, PostgreSQL Aurora, AWS Kinesis, Apache Kafka, Apache Flink, AWS S3, Iceberg, AWS Glue, Delta Lake, Terraform, CDKTF, AWS CDK, Apache Spark, Ray, GCP PubSub, Lambda, Prefect.
Responsibilities
Define long-term data architecture vision and roadmap; lead design of fault-tolerant streaming and batch infrastructure; direct database architecture and performance tuning; set technical standards for pipelines and data quality; mentor senior and mid-level engineers; evaluate and integrate modern data technologies.
Seniority
Staff, hands-on IC with technical leadership