Senior ML Engineer
Core
Own the productionization layer of Invoca's ML stack, including model serving, inference optimization, fine-tuning, and the APIs/pipelines powering the Context Engine and agentic AI workflows.
Role type
Senior ML Engineer (MLOps & Inference Infrastructure)
Builds
Production-grade ML infrastructure, inference APIs, and fine-tuned SLM/LLM models for conversation intelligence.
Domain
Conversational AI, NLP, MLOps
Deliverable
production ML models | product features | infrastructure
Required skills
MLOps (CI/CD, model versioning, monitoring), Python, PyTorch, HuggingFace Transformers, SLM/LLM fine-tuning (LoRA, QLoRA, PEFT), inference optimization (quantization, batching), Triton Inference Server, Baseten, Kubernetes, GPU infrastructure, API development, transformer-based NLP model deployment
Preferred skills
RLHF, preference training, SageMaker, Vertex AI, vLLM, TGI, Braintrust, MLflow
Technologies
Triton, Baseten, Kubernetes, PyTorch, HuggingFace, LoRA, QLoRA, PEFT, vLLM, TGI, SageMaker, Vertex AI, Braintrust, MLflow
Responsibilities
Architect and maintain CI/CD pipelines for ML artifacts; design and optimize SLM/LLM deployment for low latency and high throughput; apply parameter-efficient fine-tuning methods to adapt transformer models; contribute to model training infrastructure and data pipelines; build production-grade APIs exposing ML models to downstream consumers.
Seniority
Senior, hands-on IC