CareerPlanSign in
🌐 Remote💼 Full-time💰 $152,000–$152,000🗓 2026-06-25

Core

Own the productionization layer of Invoca's ML stack, including model serving, inference optimization, fine-tuning, and the APIs/pipelines powering the Context Engine and agentic AI workflows.

Role type

Senior ML Engineer (MLOps & Inference Infrastructure)

Builds

Production-grade ML infrastructure, inference APIs, and fine-tuned SLM/LLM models for conversation intelligence.

Domain

Conversational AI, NLP, MLOps

Deliverable

production ML models | product features | infrastructure

Required skills

MLOps (CI/CD, model versioning, monitoring), Python, PyTorch, HuggingFace Transformers, SLM/LLM fine-tuning (LoRA, QLoRA, PEFT), inference optimization (quantization, batching), Triton Inference Server, Baseten, Kubernetes, GPU infrastructure, API development, transformer-based NLP model deployment

Preferred skills

RLHF, preference training, SageMaker, Vertex AI, vLLM, TGI, Braintrust, MLflow

Technologies

Triton, Baseten, Kubernetes, PyTorch, HuggingFace, LoRA, QLoRA, PEFT, vLLM, TGI, SageMaker, Vertex AI, Braintrust, MLflow

Responsibilities

Architect and maintain CI/CD pipelines for ML artifacts; design and optimize SLM/LLM deployment for low latency and high throughput; apply parameter-efficient fine-tuning methods to adapt transformer models; contribute to model training infrastructure and data pipelines; build production-grade APIs exposing ML models to downstream consumers.

Seniority

Senior, hands-on IC

Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.