AI Developer
Core
Design, train, optimize, and deploy LLM models in on-prem and offline environments to build context-aware AI systems and end-to-end LLM pipelines.
Role type
Senior IC machine-learning engineer (LLM & RAG)
Builds
End-to-end LLM pipelines, RAG systems, and offline inference frameworks
Domain
Enterprise software, machine learning, large language models
Deliverable
production ML models
Required skills
Python, PyTorch, supervised fine-tuning, RAG pipeline design, model quantization, vector databases, Docker, Git, Azure DevOps
Preferred skills
LoRA/Q-LoRA, MCP integrations, agentic workflows, GPU optimization, air-gapped environment deployment
Technologies
LLaMA, Mistral, Qwen, Hugging Face Transformers, FAISS, Chroma, Weaviate, pgvector, vLLM, TGI, Ollama, CUDA, Postgres, MySQL
Responsibilities
Train and fine-tune LLMs using supervised fine-tuning (SFT); Design and implement end-to-end Retrieval-Augmented Generation (RAG) pipelines; Deploy and maintain models fully offline and in air-gapped environments; Build backend services in Python for ML training and inference workflows; Optimize retrieval strategies including hybrid search and re-ranking; Integrate MCP (Model Context Protocol) servers to connect LLMs with external tools and APIs
Seniority
Senior, hands-on IC