AI Engineer
Core
Designing and building production AI features for a self-hosted chat interface, including retrieval pipelines, MCP server integrations, and model routing.
Role type
Senior AI Engineer (LLM Systems & RAG)
Builds
Self-hosted chat interface, RAG pipelines, MCP server tools, and AI integrations for internal systems.
Domain
Enterprise AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG)
Deliverable
production ML models | product features
Required skills
Software engineering, Kubernetes, RAG pipeline design, vector databases, embeddings, LLM gateways (LiteLLM/OpenRouter), self-hosted model serving (vLLM/TGI/Ollama), prompt engineering, system monitoring and guardrails.
Preferred skills
MLOps/Platform experience, MCP server architecture, agent frameworks, LLM observability tools (LangSmith/Langfuse).
Technologies
Kubernetes, LiteLLM, OpenRouter, vLLM, TGI, Ollama, Open WebUI, vector stores, MCP.
Responsibilities
Maintain self-hosted chat interface and model connections; build RAG pipelines (ingestion, chunking, retrieval, reranking); integrate LLM gateways and handle routing; configure MCP server and tools; write and iterate on prompts and evaluations; monitor logging, tracing, and guardrails; ship end-to-end features including API and rollout.
Seniority
Senior, hands-on IC