Infrastructure Engineer
Core
Own and operate internal systems, infrastructure, and the AI Gateway product to power Shakudo's data and AI operating system at scale.
Role type
Senior Infrastructure Engineer (AI/ML & Cloud)
Builds
Internal services, DGX machine clusters, Kubernetes clusters, CI/CD pipelines, and the customer-facing AI Gateway product.
Domain
AI/ML infrastructure, High-Performance Computing (HPC), Cloud Operations
Deliverable
production ML models | infrastructure
Required skills
Kubernetes cluster operation, DevOps, bare-metal server operations, security hardening, observability, reliability engineering, CI/CD pipeline creation, LLM hosting and inference serving
Preferred skills
Rust programming, AI/ML infrastructure experience
Technologies
Kubernetes, DGX machines, physical servers, CI/CD systems
Responsibilities
Maintain and operate internal services including proprietary applications and ETL pipelines; Maintain and operate DGX machines hosting LLMs; Maintain and operate physical servers for Kubernetes clusters and ensure uptime; Create CI/CD pipelines for internal deployments; Maintain and operate the AI Gateway product for customers and contribute to its roadmap.
Seniority
Senior, hands-on IC