Senior Solution Architect, AI Compute Engineer - NVIS
Core
Deploy, manage, and maintain large-scale AI/HPC infrastructure in Linux environments for customers, acting as the domain expert during planning and implementation.
Role type
Senior IC AI/HPC Infrastructure Engineer
Builds
Large-scale AI/HPC systems and GPU clusters for academic and commercial customers
Domain
High-Performance Computing (HPC) and AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux System Administration, Cluster management, Scripting, Networking (advanced routing/monitoring), Troubleshooting, Documentation
Preferred skills
MPI programming, NCCL optimization, High-speed network deployment (InfiniBand/Ethernet), Automation tools (Ansible/Salt/Puppet), Kubernetes orchestration
Technologies
Linux, SLURM, LSF, UGE, MPI, NCCL, InfiniBand, Ethernet, Ansible, Salt, Puppet, Kubernetes
Responsibilities
Deploy and maintain AI/HPC infrastructure in Linux-based environments; Act as domain expert with customers during planning and implementation; Provide feedback to internal teams via bug reports and improvement suggestions; Perform knowledge transfers and documentation handovers.
Seniority
Senior, hands-on IC