Principal Software Engineer, AI Platform Engineering
Core
Architecting and governing the AI Platform's data infrastructure, ensuring tenant isolation, PII removal, and traceability for ML training signals.
Role type
Principal Software Engineer (AI Platform Engineering)
Builds
AI Data Lake, batch/streaming pipelines, schema registry, orchestration layers, feature stores, vector databases, and RAG pipelines.
Domain
Cloud Data Engineering & AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Data lake architecture, PySpark/Scala, Apache Beam/Dataflow, Schema registry (Avro/Protobuf), Orchestration (Flyte/Kubeflow), Multi-tenant isolation, Feature stores (Feast), Vector databases (Pgvector/Qdrant), RAG pipelines, gRPC/HTTPS APIs
Preferred skills
Differential privacy, Open source contributions, IAM/governance data, Iceberg/Delta Lake at petabyte scale
Technologies
GCS, Dataproc, Iceberg, Apache Beam, Dataflow, Flyte, Kubeflow, Feast, Redis, Cloud SQL, Qdrant, GKE, Great Expectations, dbt
Responsibilities
Define architectural standards for training data flow and governance; design and operate production data lakes with tiered retention and access control; build and maintain batch and streaming pipelines for feature backfills and CDC ingestion; manage schema evolution and compatibility; operate feature stores and vector databases for embedding storage and RAG; implement data quality gates and PII assertion logic.
Seniority
Principal, hands-on IC with platform-wide impact