Technical Success Engineer
Core
Own the end-to-end deployment of AI cloud infrastructure for strategic enterprise customers, validating builds and resolving technical blockers to transition from contract to live production.
Role type
Senior Technical Success Engineer (AI Infrastructure)
Builds
Production AI cloud environments (GPU/HPC clusters, Kubernetes, Linux)
Domain
AI Cloud Infrastructure / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
GPU/HPC infrastructure, cloud platforms, Kubernetes, large-scale Linux systems, troubleshooting networking/storage/compute, cross-functional coordination, technical documentation
Preferred skills
large-scale GPU cluster deployments, project tracking tools, runbook/checklist creation
Technologies
Kubernetes, Linux, GPU clusters
Responsibilities
Validate configuration, connectivity, storage, and compute against contractual promises; troubleshoot complex technical issues with engineering teams; coordinate dependencies across Infrastructure, Engineering, Product, and Data Center teams; maintain accurate deployment status and risk reports; guide customers through onboarding to first production workload; transition customers to self-sufficiency.
Seniority
Senior, hands-on IC