AI & HPC Infrastructure Engineer
Core
Design, build, and operate AI infrastructure and accelerated computing solutions for enterprise clients, enabling large-scale model training, inference, and agentic AI workloads.
Role type
Senior Infrastructure Engineer (AI & HPC)
Builds
Scalable AI infrastructure services including Bare-Metal-aaS, GPUaaS, AIaaS, and agentic AI frameworks for mission-critical clients.
Domain
AI Infrastructure, High-Performance Computing (HPC), Cloud & On-Premises Hybrid Environments
Deliverable
production ML models | infrastructure
Required skills
Cluster management, workload scheduling, Kubernetes orchestration, accelerated computing platforms (GPU/DPU/LPU), high-speed interconnects, AI storage architectures, infrastructure automation, observability, NVIDIA platform tools, LLM inference engines, production serving frameworks
Preferred skills
MLOps/LLMOps frameworks, API development, agentic AI infrastructure design, MCP server integration, specialty cloud platform deployment, machine learning frameworks
Technologies
Kubernetes, Slurm, Run:ai, AWS, Azure, GCP, VMware, Nutanix, Python, Terraform, Ansible, NVIDIA Base Command Manager, NGC, NCCL, NVLink, CUDA, TensorRT-LLM, vLLM, SGLang, Triton Inference Server, MLPerf, InfiniBand, SONiC, NVMe, Weka, DDN, TensorFlow, PyTorch, JAX
Responsibilities
Design and implement AI infrastructure and accelerated computing solutions; Deploy and manage XPU-based clusters across bare-metal and containerized environments; Integrate AI platforms with IT systems and security frameworks; Architect agentic AI infrastructure with observability and policy controls; Build MCP servers and adapters for infrastructure monitoring; Develop architecture diagrams and operational runbooks; Provide technical guidance and optimization for AI workloads
Seniority
Senior, hands-on IC