CareerPlanSign in

ML / RAG / Inference Engineer

Onsite or remote • Noida+1💼 Full-time🗓 2026-06-24 → 2026-09-22

Core

Build the AI orchestration layer for Federis to govern retrieval, prompts, model access, inference routing, evaluation, and observability across replaceable external AI and vector systems.

Role type

Senior IC ML / RAG / Inference Engineer

Builds

AI orchestration layer, provider adapters, evaluation pipelines, and scalable data/inference paths

Domain

Enterprise AI, RAG systems, vector search, model serving

Deliverable

production ML models

Required skills

RAG workflows, retrieval configuration, prompt/version governance, inference routing, evaluation hooks, model usage telemetry, provider adapters, quality/latency/cost/safety evaluation pipelines, policy enforcement, scalable data/inference path design

Preferred skills

vLLM, Triton, Milvus, Qdrant, Weaviate, OpenSearch, Apache Solr, customer-hosted LLM stacks, AI evaluation harnesses, guardrails, prompt governance, red-team tests, model risk workflows, regulated enterprise AI, data governance, knowledge management, secure document intelligence

Technologies

Python, TypeScript, Go, Java, vector databases, model serving systems, embedding services, customer-managed inference endpoints

Responsibilities

Implement RAG workflows and retrieval configuration; Build provider adapters for vector databases and model serving systems; Create evaluation pipelines for quality, latency, cost, grounding, and safety; Partner with security teams on policy enforcement; Design scalable data and inference paths for SaaS, private cloud, and air-gapped environments; Document model and retrieval behavior for solution engineers and auditors

Seniority

Senior, hands-on IC

Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.