Software Engineer, Data Foundations
Core
Build and scale the end-to-end data ingestion and management layer powering Glean's Search, AI Assistant, and Agent products across thousands of enterprise apps and billions of documents.
Role type
Senior backend/data infrastructure engineer
Builds
Enterprise connectors, data pipelines, and permission-aware document representations for AI reasoning
Domain
Enterprise SaaS integration, data infrastructure, AI/ML data foundations
Deliverable
production ML models | infrastructure
Required skills
distributed systems, data pipelines, queues, large-scale storage (SQL/NoSQL), strict consistency, permission modeling, SLOs, error budgets, exactly-once processing, backpressure, retries, observability
Preferred skills
enterprise connectors, search/indexing, information retrieval, security-sensitive systems, LLMs and AI tools
Technologies
Java, Go, C++, Python, Google Workspace, Microsoft 365, Slack, Salesforce, Jira, ServiceNow, GitHub
Responsibilities
Build and scale connectors to SaaS and on-prem systems; Handle full syncs, low-latency incremental updates, rate-limiting, and complex authentication flows; Transform raw, unstructured enterprise content into rich, structured, permission-aware representations; Design document schemas and enrichment pipelines; Own end-to-end correctness, freshness, and performance for petabyte-scale data flows; Solve hard problems in ordering, idempotency, exactly-once processing, backpressure, and retries; Preserve fine-grained ACLs, deletions, and sensitivity constraints; Partner with Search Serving, Product, Platforms, and Security teams; Improve observability, alerting, and automation
Seniority
Senior, hands-on IC