CareerPlanGet AI match score →

ML / RAG / Inference Engineer

Onsite or remote • Noida+1💼 Full-time🗓 2026-06-24

Core

Build the AI orchestration layer for Federis to govern retrieval, prompts, model access, inference routing, evaluation, and observability across replaceable external AI and vector systems.

Role type

Senior IC ML / RAG / Inference Engineer

Builds

AI orchestration layer, provider adapters, evaluation pipelines, and scalable data/inference paths

Domain

Enterprise AI, RAG systems, vector search, model serving

Deliverable

production ML models

Required skills

RAG workflows, retrieval configuration, prompt/version governance, inference routing, evaluation hooks, model usage telemetry, provider adapters, quality/latency/cost/safety evaluation pipelines, policy enforcement, scalable data/inference path design

Preferred skills

vLLM, Triton, Milvus, Qdrant, Weaviate, OpenSearch, Apache Solr, customer-hosted LLM stacks, AI evaluation harnesses, guardrails, prompt governance, red-team tests, model risk workflows, regulated enterprise AI, data governance, knowledge management, secure document intelligence

Technologies

Python, TypeScript, Go, Java, vector databases, model serving systems, embedding services, customer-managed inference endpoints

Responsibilities

Implement RAG workflows and retrieval configuration; Build provider adapters for vector databases and model serving systems; Create evaluation pipelines for quality, latency, cost, grounding, and safety; Partner with security teams on policy enforcement; Design scalable data and inference paths for SaaS, private cloud, and air-gapped environments; Document model and retrieval behavior for solution engineers and auditors

Seniority

Senior, hands-on IC

Rewrite
## About the role Build the AI orchestration layer for Federis so customers can govern retrieval, prompts, model access, inference routing, evaluation, and observability across replaceable external AI and vector systems. ## Key Responsibilities - Implement RAG workflows, retrieval configuration, prompt/version governance, inference routing, evaluation hooks, and model usage telemetry. - Build provider adapters for vector databases, model serving systems, embedding services, and customer-managed inference endpoints. - Create quality, latency, cost, grounding, and safety evaluation pipelines that support enterprise release gates and customer reporting. - Partner with security and product teams on policy enforcement for data access, prompt execution, model selection, and audit capture. - Design scalable data and inference paths that work across SaaS, private cloud, and air-gapped customer environments. - Document model and retrieval behavior clearly enough for solution engineers, customers, and auditors to understand. ## Required Experience - 4+ years in ML engineering, search, data platforms, RAG systems, model serving, or AI product engineering. - Strong software engineering ability in Python, TypeScript, Go, Java, or similar production languages. - Practical experience with embeddings, vector search, ranking, chunking, retrieval evaluation, prompt/version management, and model APIs. - Understanding of latency, throughput, cost, reliability, privacy, and observability tradeoffs in AI systems. - Ability to build adapter-based integrations without binding the core product to a single AI vendor or database. ## Useful Differentiators - Experience with vLLM, Triton, Milvus, Qdrant, Weaviate, OpenSearch, Apache Solr, or customer-hosted LLM stacks. - Experience creating AI evaluation harnesses, guardrails, prompt governance, red-team tests, or model risk workflows. - Prior work in regulated enterprise AI, data governance, knowledge management, or secure document intelligence. ## Success Scorecard - AI integrations are replaceable and measurable across quality, latency, cost, safety, and audit dimensions. - RAG workflows produce reproducible evidence for what data was used, what model was called, and what policy applied. - Customer deployments can use their preferred inference and vector systems without core rewrites. - Evaluation results inform roadmap priorities and customer success plans.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗