HPC Specialist
Core
Design, deploy, and optimize GPU infrastructure for large-scale LLM inference and ML workloads within a systematic trading environment.
Role type
Senior HPC Specialist (GPU Infrastructure & Model Serving)
Builds
GPU server fleets, distributed model serving solutions, and Kubernetes clusters for AI/ML workloads.
Domain
High-Performance Computing, AI/ML Infrastructure, Systematic Trading
Deliverable
infrastructure
Required skills
GPU infrastructure management, model serving frameworks (vLLM, SGLang), Linux systems administration, Kubernetes orchestration, distributed systems architecture, network configuration, storage optimization, Python scripting, Bash scripting, infrastructure as code (Ansible, Terraform), monitoring and observability.
Preferred skills
Experience optimizing deep learning workloads, troubleshooting performance bottlenecks across hardware and application layers, research on emerging GPU technologies.
Technologies
vLLM, SGLang, Kubernetes, Prometheus, Grafana, Ansible, Terraform
Responsibilities
Deploy and maintain GPU infrastructure for LLM inference; Architect distributed serving solutions for multi-node deployments; Manage GPU-enabled Kubernetes clusters; Configure network infrastructure for GPU clusters; Implement storage solutions for model weights; Troubleshoot performance bottlenecks; Research emerging GPU technologies; Collaborate with ML engineers on inference acceleration; Drive reliability improvements through monitoring and capacity planning.
Seniority
Senior, hands-on IC