Manager, Solutions Architecture - Continuous Bringup and Optimization
Core
Lead a team to consult, optimize, and improve the resiliency of customer AI factory infrastructures, ensuring high service quality and operational perfection for large-scale AI/HPC systems.
Role type
Manager, Solutions Architecture (Infrastructure & Operations)
Builds
Large-scale AI/HPC systems, data center environments, and optimized GPU-accelerated workflows.
Domain
AI/HPC, Data Center Operations, Networking, System Design
Deliverable
production ML models | infrastructure
Required skills
Team leadership, infrastructure analysis and tuning, data center architecture, GPU/CPU/networking expertise, Linux environments, root cause analysis, continuous improvement methodologies
Preferred skills
AI infrastructure workflows, MLOps/DevOps tools, containerization (Docker, Kubernetes), large-scale system deployments, data center safety and security protocols
Technologies
NVIDIA GPU/CPU, Networking topologies, Linux, Docker, Kubernetes
Responsibilities
Lead team consulting and optimization of customer AI factory infrastructures; Drive hands-on analysis and tuning of complex GPU-accelerated systems; Align infrastructure strategies with business goals via internal collaboration; Act as technical authority on NVIDIA technologies for customer reviews; Establish optimization and monitoring methodologies; Participate in customer-facing engagements including roadmap sessions and incident retrospectives
Seniority
Manager, hands-on IC with team leadership