ML / RAG / Inference Engineer
Core
Build the AI orchestration layer for Federis to govern retrieval, prompts, model access, inference routing, evaluation, and observability across replaceable external AI and vector systems.
Role type
Senior IC ML / RAG / Inference Engineer
Builds
AI orchestration layer, provider adapters, evaluation pipelines, and scalable data/inference paths
Domain
Enterprise AI, RAG systems, vector search, model serving
Deliverable
production ML models
Required skills
RAG workflows, retrieval configuration, prompt/version governance, inference routing, evaluation hooks, model usage telemetry, provider adapters, quality/latency/cost/safety evaluation pipelines, policy enforcement, scalable data/inference path design
Preferred skills
vLLM, Triton, Milvus, Qdrant, Weaviate, OpenSearch, Apache Solr, customer-hosted LLM stacks, AI evaluation harnesses, guardrails, prompt governance, red-team tests, model risk workflows, regulated enterprise AI, data governance, knowledge management, secure document intelligence
Technologies
Python, TypeScript, Go, Java, vector databases, model serving systems, embedding services, customer-managed inference endpoints
Responsibilities
Implement RAG workflows and retrieval configuration; Build provider adapters for vector databases and model serving systems; Create evaluation pipelines for quality, latency, cost, grounding, and safety; Partner with security teams on policy enforcement; Design scalable data and inference paths for SaaS, private cloud, and air-gapped environments; Document model and retrieval behavior for solution engineers and auditors
Seniority
Senior, hands-on IC