Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp
Core
Principal Technical Account Manager guiding enterprise customers through AI/ML and HPC implementation journeys, focusing on distributed training, GPU/Trainium cluster architecture, and production-grade inference.
Role type
Principal Technical Account Manager (AI/ML HPC Specialist)
Builds
Production-grade AI/ML training and inference solutions, distributed training clusters, and MLOps pipelines for enterprise customers.
Domain
Cloud Computing, Artificial Intelligence, Machine Learning, High-Performance Computing (HPC)
Deliverable
production ML models | infrastructure
Required skills
Distributed training architecture, GPU/Trainium cluster management, NCCL communication tuning, Slurm job scheduling, Deep Learning frameworks (PyTorch, DeepSpeed, Megatron-LM), AWS Parallel Computing Service (PCS), SageMaker HyperPod, MLOps pipelines, LLM deployment, HPC-to-AI convergence patterns
Preferred skills
Experience with frontier models (multi-trillion parameters), AWS Neuron SDK, EC2 Capacity Blocks, cluster observability, automated failure recovery
Technologies
AWS ParallelCluster, P6/P6e, G7/G7e instances, Trn3 UltraServers, EFA SRD, Slurm, PyTorch FSDP, DeepSpeed, Megatron-LM, SageMaker HyperPod, Amazon FSx for Lustre, DLAMIs, AWS Step Functions, AWS Batch, Elastic Fabric Adapter
Responsibilities
Lead technical deep-dives and performance optimization for enterprise AI/ML workloads; Design and implement production-grade AI/ML training and inference solutions; Support customers in implementing business-critical HPC capabilities including LLMs and PINNs; Partner with service teams to enhance model training throughput and optimize GPU utilization; Serve as trusted advisor for AI/ML infrastructure decisions spanning compute, networking, and storage.
Seniority
Principal, hands-on IC with strategic advisory