Software Engineer - LLM Infrastructure
Core
Develop, maintain, and scale a self-hosted, custom-tuned Large Language Model (LLM) to provide context-aware insights into complex system interactions and operational dynamics for distributed infrastructure.
Role type
Senior IC Machine Learning Infrastructure Engineer (LLM)
Builds
Self-hosted LLM inference engine and RAG information retrieval systems for SwoopOS
Domain
Defense technology, distributed robotic systems, and critical infrastructure
Deliverable
production ML models
Required skills
Scalable LLM infrastructure, NVIDIA drivers and Kubernetes Container Toolkit, RAG system design, time-series anomaly detection, PyTorch, Kubernetes production environment, Python, data structures and data models
Preferred skills
On-premise/self-hosted AI, multi-GPU performance optimization, vLLM inference engines
Technologies
PyTorch, Kubernetes, NVIDIA, vLLM
Responsibilities
Develop and maintain Swoop's LLM offering, expand LLM capabilities for tool calls and data searching, monitor resource usage and optimize inference efficiency, maintain and optimize inference engine architecture, tune data storage configurations for scale and streaming availability, ensure strong service availability for the Kubernetes cluster runtime
Seniority
Senior, hands-on IC