Customer Support Engineer (GPU Cluster), India
Core
First-line technical support for customers building training, fine-tuning, and inference solutions on Kubernetes GPU clusters.
Role type
Senior Customer Support Engineer (GPU Infrastructure)
Builds
Production GPU cluster infrastructure for AI workloads
Domain
AI/ML Infrastructure, High-Performance Computing (HPC)
Deliverable
client delivery
Required skills
Kubernetes, GPU technologies, HPC environments, infrastructure services (SLURM), infrastructure as code (Ansible), container infrastructure, scripting/programming languages, cluster administration, complex technical troubleshooting
Preferred skills
AI/ML domain knowledge, cross-functional collaboration, documentation, pattern recognition from support cases
Technologies
Kubernetes, SLURM, Ansible, NFS
Responsibilities
Resolve complex technical challenges involving Kubernetes GPU clusters; act as last line of technical defense before escalation; collaborate with Engineering, Research, and Product teams; transform customer insights into roadmap actions; maintain detailed documentation of system configurations and troubleshooting guides; provide flexible support coverage during holidays, nights, and weekends
Seniority
Mid-Senior, hands-on IC