Senior Hardware engineer (R&D / GPU / AI)
Core
Design, deploy, and maintain high-performance cloud systems optimized for AI workloads, including troubleshooting complex hardware, software, and networking issues in data center environments.
Role type
Senior System Engineer (Hardware R&D)
Builds
High-performance cloud systems for AI workloads
Domain
Cloud infrastructure / AI compute / Hardware systems
Deliverable
infrastructure
Required skills
Server architecture, GPU technologies, Networking (InfiniBand, NVLink, PCIe), Linux systems, Python, Bash, Root cause analysis, Performance optimization, Electronics modification (soldering, wiring)
Preferred skills
Linux kernel debugging, Electronic measurement equipment usage
Technologies
InfiniBand, NVLink, PCIe, Linux, Python, Bash
Responsibilities
Participate in the design, deployment, and maintenance of high-performance cloud systems optimized for AI workloads; Arrange and perform hardware R&D tests and experiments on-site in data center environments; Troubleshoot and resolve complex system issues related to GPUs, networking, and server infrastructure; Conduct deep investigations into hardware, software, and networking issues to ensure optimal system performance and reliability; Develop and execute test plans and methodologies for advanced GPU, InfiniBand, and compute systems to benchmark and validate performance; Monitor system performance and continuously fine-tune configurations for maximum efficiency.
Seniority
Senior, hands-on IC