AI Platform Architect
Core
Design and oversee the comprehensive infrastructure stack powering distributed AI workloads, acting as the unifying technical authority across hardware, software, compute, network, and storage for trillion-parameter LLM training and high-throughput inference.
Role type
Senior IC AI Platform Architect
Builds
Cohesive AI rack scale platforms optimized for large model training and inference
Domain
AI Infrastructure / HPC / Systems Engineering
Deliverable
production ML models
Required skills
Systems engineering, cloud architecture, HPC, distributed AI frameworks, systems interconnects, container orchestration, cross-domain leadership
Preferred skills
Rack scale GPU AI platforms experience, AI platforms integrating latest networking/cooling/GPU tech, Python/JSON scripting
Technologies
Kubernetes, Slurm, PyTorch, DeepSpeed, PCIe Gen 5/6, NVMe, RDMA (RoCEv2/InfiniBand), ARM mesh interconnects (RNI, HNF, SNF)
Responsibilities
Define holistic architecture for highly clustered AI environments ensuring zero-bottleneck data flow; Influence strategy for AI workload scheduling and orchestration at massive scale; Profile and eliminate system-level bottlenecks across the entire AI pipeline; Work closely with software, firmware, and OS engineering to influence platform design; Drive the 3-to-5-year technical vision for the AI platform and influence silicon design requirements
Seniority
Senior, hands-on IC