Staff ML/LLM Ops Engineer
Core
Building a standardized self-serve MLOps/LLMOps platform to manage the lifecycle of computer-vision, LLM, VLM, and agentic workloads, ensuring safe, versioned, and observable model deployment from research to production.
Role type
Staff ML/LLM Ops Engineer (Senior IC with technical leadership)
Builds
A unified model serving and orchestration platform for vision and generative AI workloads
Domain
Intelligent site technology, AI infrastructure, Computer Vision, Generative AI
Deliverable
production ML models
Required skills
MLOps platform design and operation, LLM/VLM operationalization, CI/CD for models, API contract design, technical mentorship, polyglot service architecture, production observability, model registry management, drift detection, guardrails implementation
Preferred skills
Computer vision inference at scale, cloud-native infrastructure (Kubernetes/Argo), building ML platforms from scratch, edge deployment (NVIDIA Jetson), agentic tooling (LangGraph/MCP/vector databases)
Technologies
Kubernetes, Argo, NVIDIA Jetson, LangGraph, MCP, vector databases
Responsibilities
Own the end-to-end model lifecycle including packaging, CI/CD, serving, versioning, and monitoring; Implement LLM-specific guardrails, evaluation suites, and cost/latency observability; Design self-serve deployment pipelines and API boundaries; Mentor engineers on productionization standards
Seniority
Staff, hands-on IC with technical leadership