CareerPlanSign in
💼 Full-time🗓 2026-06-25

Core

Develop, maintain, and scale a self-hosted, custom-tuned Large Language Model (LLM) to provide context-aware insights into complex system interactions and operational dynamics for distributed infrastructure.

Role type

Senior IC Machine Learning Infrastructure Engineer (LLM)

Builds

Self-hosted LLM inference engine and RAG information retrieval systems for SwoopOS

Domain

Defense technology, distributed robotic systems, and critical infrastructure

Deliverable

production ML models

Required skills

Scalable LLM infrastructure, NVIDIA drivers and Kubernetes Container Toolkit, RAG system design, time-series anomaly detection, PyTorch, Kubernetes production environment, Python, data structures and data models

Preferred skills

On-premise/self-hosted AI, multi-GPU performance optimization, vLLM inference engines

Technologies

PyTorch, Kubernetes, NVIDIA, vLLM

Responsibilities

Develop and maintain Swoop's LLM offering, expand LLM capabilities for tool calls and data searching, monitor resource usage and optimize inference efficiency, maintain and optimize inference engine architecture, tune data storage configurations for scale and streaming availability, ensure strong service availability for the Kubernetes cluster runtime

Seniority

Senior, hands-on IC

Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.