AI Compute Engineer
Core
Design and validate high-performance bare-metal GPU infrastructure for large-scale AI training and inference, ensuring operational readiness across data centers. (via careerplan.io/jobs/f07fd68b-abe2-4fb4-a4c7-d0e432ea9b98-ai-compute-engineer-at-breakmark)
Role type
Senior IC AI Compute Architect
Builds
Production-ready GPU clusters and infrastructure platforms
Domain
AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
NVIDIA HGX platform architecture, GPU cluster qualification and burn-in, fleet-level failure analysis, Linux systems administration, NVIDIA software stack (drivers, CUDA, NCCL, DCGM), root-cause analysis of hardware/software/firmware issues, distributed workload validation, cross-domain integration (network/storage/facilities), automation requirement definition
Preferred skills
NVIDIA B200/B300 experience, InfiniBand/ConnectX networking, large-scale NCCL qualification, PyTorch/NeMo/Megatron/vLLM/TensorRT-LLM, GPU schedulers (Slurm/Kubernetes/Run:ai), shared storage validation (WEKA/VAST), infrastructure automation platforms, Python/Bash scripting
Technologies
NVIDIA HGX, H100/H200/B200/B300, NVLink, NVSwitch, PCIe, RDMA, CUDA, NCCL, DCGM, Slurm, Kubernetes, Run:ai, WEKA, VAST, InfiniBand, ConnectX, PyTorch, NeMo, Megatron, vLLM, TensorRT-LLM
Responsibilities
Define lifecycle standards for GPU platforms and validated baselines for firmware/BIOS/OS/drivers; Design qualification strategies for new clusters and analyze fleet-wide performance distributions; Diagnose health and performance issues using NVIDIA telemetry and lead root-cause analysis; Validate intra-node and multi-node GPU communication with real workloads; Integrate across network, storage, DevOps, and facilities domains; Own operational readiness gates and guide automation teams in building reproducible workflows
Seniority
Senior, hands-on IC