Service Engineer II - CTJ - Poly
Core
Support production Datacenter IT process escalations, perform deep-dive root cause analysis, and develop automation/tools to improve analytics and reduce manual mitigations.
Role type
Senior IC Service Engineer (Datacenter Infrastructure & AI Systems)
Builds
Production Datacenter operations, automated diagnostic tools, and standard operating procedures for high-performance computing environments.
Domain
Datacenter operations, AI/ML computing infrastructure, high-speed interconnects, and industrial controls.
Deliverable
production ML models | infrastructure
Required skills
Root cause analysis, automation development, GPU debugging (NVIDIA tools), distributed systems experience, cloud infrastructure management, programming (C++, Python, Java, etc.), datacenter environment concepts (power/cooling), engineering methodologies.
Preferred skills
Experience with InfiniBand, RDMA, FPGAs, ASICs, deep learning frameworks, agile development.
Technologies
NVIDIA SMI, NVIDIA Field Diagnostics, NVIDIA Data Center GPU Manager, InfiniBand, RDMA, C++, C#, Java, PowerShell, Python, JavaScript.
Responsibilities
Mitigate production incidents rapidly, define and drive repair actions for root causes, create new tools integrating data from various systems, develop standard procedures and technical service guides, report on high-impact issues in engineering forums.
Seniority
Senior, hands-on IC