ML Engineering Lead (LLM Ops)
Core
Own the operational lifecycle of Neko's LLM, GenAI, and RAG-based systems, building a production-grade platform for clinical ML workflows on proprietary sensor/device data.
Role type
ML Engineering Lead (LLM Ops)
Builds
Production LLM and agent pipelines, evaluation suites, and serving strategy frameworks for clinical use cases.
Domain
Healthcare / Medical Devices / Generative AI
Deliverable
production ML models
Required skills
MLOps lifecycle management, Python, LLM application development, prompt engineering, RAG, agentic workflows, PyTorch, distributed systems, ML orchestration, vector databases, embedding models, human-feedback loop integration
Preferred skills
Agentic/AI-assisted coding workflows, Kubernetes, Terraform, LLM evaluation and observability practices, Databricks MLflow 3 for GenAI
Technologies
MLflow, Unity Catalog, LangChain, LangGraph, PyTorch, Databricks, Kubernetes, Terraform
Responsibilities
Stand up MLflow Tracing observability across production LLM pipelines; Build evaluation suites with LLM judges and human-feedback loops; Ship RAG or agentic pipelines with versioning; Produce cost, latency, and GPU capacity frameworks; Integrate LLM Ops tightly with the existing MLOps team
Seniority
Lead, hands-on IC with strategic ownership (via careerplan.io/jobs/b8648fba-ee52-47c2-8589-4b90da34c2b1-ml-engineering-lead-llm-ops-at-neko-health)