Forward Deployed Engineer (Generative AI)
Core
Deploy, integrate, and scale enterprise Generative AI solutions and Large Language Models (LLMs) within customer cloud environments, bridging AI research with production-grade infrastructure.
Role type
Senior IC Forward Deployed Engineer (Generative AI)
Builds
Production-grade LLM orchestration frameworks, RAG systems, and scalable AI infrastructure on Google Cloud Platform.
Domain
Generative AI, Cloud Infrastructure, Enterprise Consulting
Deliverable
production ML models
Required skills
LLM orchestration (LangChain, LlamaIndex, AutoGen), Deep learning frameworks (PyTorch, Hugging Face), Vector databases (Vertex AI Vector Search, Milvus, Pinecone, pgvector), Model serving (vLLM, TGI, Triton), Kubernetes (GKE), Infrastructure as Code (Terraform), Python/Go
Preferred skills
None stated
Technologies
Google Cloud Platform, Vertex AI, GKE, Terraform, LangChain, LlamaIndex, AutoGen, PyTorch, Hugging Face, vLLM, TGI, Triton, Milvus, Pinecone, pgvector
Responsibilities
Deploy and optimize large-scale Gen AI models and LLM frameworks; Architect scalable infrastructure for GPU/TPU workloads; Design high-throughput data ingestion and Vector Database architectures for RAG; Act as primary technical consultant for AI safety and cost optimization; Feed deployment insights back to core research teams.
Seniority
Senior, hands-on IC