Data Infra Engineer
Core
Design, develop, and maintain high-throughput distributed backend components and services for the Commercialization Data Lake, building scalable ingestion pipelines for petabytes of telemetry and RPC logs.
Role type
Senior IC data infrastructure engineer
Builds
Scalable data ingestion pipelines, canonical base tables, and self-service Data Pipeline Services
Domain
Data engineering / Distributed systems
Deliverable
production ML models | product features | infrastructure
Required skills
Distributed systems architecture, Database management, High-throughput batch and streaming pipeline design, C++, SQL, Regression testing, Canary/progressive rollout strategies, DAG architecture
Preferred skills
Spark, Flink, Beam, Hadoop, Flume, Data warehouse or data lake systems, Data governance and privacy capabilities (encryption, GDPR, RBAC/ACLs), Google-internal pipeline stacks (Flume, SQLP, Rapid, CAS, F1)
Responsibilities
Design and maintain high-throughput distributed backend components and services, Build scalable ingestion pipelines for raw system telemetry and RPC logs, Develop standardized Data Pipeline Services and self-service gateways, Implement compliance, privacy, security, and identity infrastructure, Establish regression testing and operational standards for pipeline reliability
Seniority
Senior, hands-on IC