CareerPlanSign in

AI Compute Engineer

ColombiaFull-time2026-10-05 → 2026-10-07

Core

Design and validate high-performance bare-metal GPU infrastructure for large-scale AI training and inference, ensuring operational readiness across data centers. (via careerplan.io/jobs/f07fd68b-abe2-4fb4-a4c7-d0e432ea9b98-ai-compute-engineer-at-breakmark)

Role type

Senior IC AI Compute Architect

Builds

Production-ready GPU clusters and infrastructure platforms

Domain

AI Infrastructure / High-Performance Computing

Deliverable

production ML models

Required skills

NVIDIA HGX platform architecture, GPU cluster qualification and burn-in, fleet-level failure analysis, Linux systems administration, NVIDIA software stack (drivers, CUDA, NCCL, DCGM), root-cause analysis of hardware/software/firmware issues, distributed workload validation, cross-domain integration (network/storage/facilities), automation requirement definition

Preferred skills

NVIDIA B200/B300 experience, InfiniBand/ConnectX networking, large-scale NCCL qualification, PyTorch/NeMo/Megatron/vLLM/TensorRT-LLM, GPU schedulers (Slurm/Kubernetes/Run:ai), shared storage validation (WEKA/VAST), infrastructure automation platforms, Python/Bash scripting

Technologies

NVIDIA HGX, H100/H200/B200/B300, NVLink, NVSwitch, PCIe, RDMA, CUDA, NCCL, DCGM, Slurm, Kubernetes, Run:ai, WEKA, VAST, InfiniBand, ConnectX, PyTorch, NeMo, Megatron, vLLM, TensorRT-LLM

Responsibilities

Define lifecycle standards for GPU platforms and validated baselines for firmware/BIOS/OS/drivers; Design qualification strategies for new clusters and analyze fleet-wide performance distributions; Diagnose health and performance issues using NVIDIA telemetry and lead root-cause analysis; Validate intra-node and multi-node GPU communication with real workloads; Integrate across network, storage, DevOps, and facilities domains; Own operational readiness gates and guide automation teams in building reproducible workflows

Seniority

Senior, hands-on IC