Customer Success Engineer (CSE), GPU Cluster
Core
Technical owner for strategic customer relationships, ensuring flawless delivery and operational health of large-scale GPU deployments across compute, networking, storage, and facilities.
Role type
Senior Customer Success Engineer (Infrastructure)
Builds
Large-scale GPU clusters and AI infrastructure for strategic enterprise customers
Domain
AI Infrastructure / High-Performance Computing (HPC)
Deliverable
production ML models | infrastructure
Required skills
GPU infrastructure management, large-scale Ethernet and InfiniBand fabric architecture, enterprise storage systems, DC operations and facilities coordination, incident management and RCA authorship, infrastructure monitoring and observability, project management for capacity expansions
Preferred skills
Python, Bash, infrastructure automation tools
Technologies
Prometheus, Grafana, InfiniBand, NVMe, parallel file systems
Responsibilities
Serve as named technical point of contact for strategic customers; lead issue lifecycle management and RCA authorship; own end-to-end RMA coordination and hardware lifecycle management; maintain deep technical expertise in GPU compute and storage stacks; own observability strategy including alert policies and dashboards; coordinate DC operations and facilities events; act as project manager for capacity expansions and node deployment
Seniority
Senior, hands-on IC