Senior AI Compute Infrastructure Engineer
Core
Building and operating GPU and accelerator infrastructure to power AI model training, inference, evaluation, and experimentation for a global crypto exchange.
Role type
Senior AI Compute Infrastructure Engineer
Builds
GPU clusters, scheduling systems, inference pipelines, and observability tooling for internal AI workloads.
Domain
Cryptocurrency exchange / AI Infrastructure / High-Performance Computing
Deliverable
infrastructure
Required skills
GPU cluster operations, distributed systems, Kubernetes, Python, ML serving frameworks (vLLM, Triton, TensorRT), performance optimization, cost management, observability, incident response
Preferred skills
Custom silicon experience (TPUs, Trainium), capacity planning, distributed training frameworks (DeepSpeed, Ray), systems programming (Rust, C++, CUDA), crypto/trading infrastructure
Technologies
Kubernetes, vLLM, Triton Inference Server, TensorRT, Ray, DeepSpeed, CUDA, Linux
Responsibilities
Operate GPU and accelerator clusters for training and inference; design infrastructure for local GPU model execution; build scheduling and orchestration systems; optimize inference pipelines for latency and cost; partner with ML teams to remove bottlenecks; build observability for GPU utilization and spend; drive reliability and incident response; evaluate new hardware and frameworks; build tooling for GPU usage visibility; contribute to long-term architecture decisions
Seniority
Senior, hands-on IC