CareerPlanGet AI match score →
💼 Full-time🗓 2026-06-25

Core

Own the end-to-end LLM pipeline from retrieval architecture to fine-tuning and production deployment for a 0-1 AI product.

Role type

Senior IC machine-learning engineer (LLM/RAG)

Builds

Retrieval-Augmented Generation systems and LLM pipelines

Domain

Artificial Intelligence / Large Language Models

Deliverable

production ML models

Required skills

LLMs (Llama, Mistral, GPT-4, Claude), LangChain, LlamaIndex, hybrid search, reranking, vector databases (Pinecone, Weaviate), LoRA/QLoRA, Hugging Face Transformers, Python, FastAPI, PyTorch, Redis, Postgres, Docker, Kubernetes, automated evaluation loops

Preferred skills

IoT/Hardware integration, real-time systems, high-performance startups, defense-related AI projects

Technologies

Llama, Mistral, GPT-4, Claude, LangChain, LlamaIndex, Pinecone, Weaviate, FastAPI, PyTorch, Redis, Postgres, Docker, Kubernetes

Responsibilities

Build and maintain the full LLM lifecycle (Retrieval, Evaluation, Fine-tuning); Design and optimize RAG systems using hybrid search; Make critical architecture decisions for scale and sub-second speed; Monitor and improve accuracy and evaluation metrics; Translate business needs into technical AI roadmaps

Seniority

Senior, hands-on IC

Rewrite
## About the role We are looking for an engineer who can own the LLM pipeline end-to-end. You will be responsible for everything from initial retrieval architecture to fine-tuning and production deployment. This is a 0-1 role where you will work directly with founders and domain experts to define the future of AI at Interexy. ## Responsibilities - Pipeline Ownership: Build and maintain the full LLM lifecycle: Retrieval → Evaluation → Fine-tuning. - RAG Architecture: Design and optimize Retrieval-Augmented Generation systems using hybrid search (BM25 + Vector). - System Scaling: Make critical architecture decisions to ensure scale, reliability, and sub-second speed. - Continuous Improvement: Monitor and improve accuracy, grounding, and evaluation metrics (Recall@k, Precision@k). - Founder Collaboration: Work in a flat structure to translate business needs into technical AI roadmaps. ## Technical Qualifications ### Must-Have: - Core AI: Deep experience with LLMs (Llama, Mistral, GPT-4, Claude) and frameworks (LangChain / LlamaIndex). - Search & Retrieval: Proficiency in hybrid search, reranking, and vector databases (Pinecone, Weaviate). - Fine-Tuning: Hands-on experience with LoRA/QLoRA and Hugging Face Transformers. - Engineering: Expert level in Python, with solid experience in FastAPI, PyTorch, Redis, and Postgres. - Infrastructure: Proficiency in Docker & Kubernetes for deploying AI systems at scale. - Evaluation: Experience building automated eval loops (groundedness, hallucination checks). ### Preferred: - Experience with IoT/Hardware integration or real-time systems. - Prior experience in high-performance startups or defense-related AI projects. - Degree in Computer Science or Applied ML from a top-tier university. ## Soft Skills (The 'Interexy' DNA) - Fast Execution: You prefer shipping code over endless meetings. - High Ownership: You don't wait for tickets; you identify problems and solve them. - Resilience: You thrive under pressure and enjoy the 0-1 phase of product development. ## What We Offer - Impact: Real ownership of the AI roadmap in an international company. - Team: Work with a tight-knit, high-output engineering team. - Growth: Flat structure with a direct line to leadership and equity opportunities.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗