Staff Software Development Engineer (Data Engineer)
Core
Architecting the data infrastructure and retrieval systems that power semantic search, graph-based reasoning, and GenAI platforms for global information discovery.
Role type
Staff Software Development Engineer (Data Engineer)
Builds
Scalable data pipelines, hybrid retrieval architectures, knowledge graphs, and automated data validation systems for AI agents.
Domain
Artificial Intelligence / Semantic Search / Knowledge Graphs
Deliverable
production ML models
Required skills
Distributed system design, Vector Databases (HNSW, IVFFlat, PQ), Graph Databases (Cypher, Gremlin), LLM orchestration (LangChain, LlamaIndex), Large-scale data processing (Spark, Flink, Kafka), Cloud-native infrastructure (AWS/GCP/Azure), Container orchestration (Kubernetes)
Preferred skills
Semantic layering, Ontology design, Information Retrieval fundamentals (BM25, TF-IDF)
Technologies
Python, Java, Go, Pinecone, Milvus, Weaviate, Qdrant, Neo4j, AWS Neptune, ArangoDB, LangChain, LlamaIndex, OpenAI, HuggingFace, Cohere, Spark, Flink, Kafka, Kubernetes, HNSW, IVFFlat, PQ
Responsibilities
Lead design of hybrid retrieval architectures combining vector similarity search with structured graph traversals; Architect scalable data pipelines for ingestion, embedding, and indexing of massive multi-modal datasets; Innovate and prototype advanced retrieval techniques including multi-stage re-ranking and graph-tooling for LLMs; Design and implement schemas for complex knowledge graphs ensuring high-performance relationship mapping; Build automated data validation and drift detection systems; Drive technical implementation of "Memory" systems for AI agents; Champion data organization standards to unify disparate sources into a coherent knowledge base; Collaborate with AI Research and Product teams to evaluate and integrate emerging database technologies.
Seniority
Staff, hands-on IC with architectural leadership