Cloud and AI System Intern
Core
Research and prototype detection/diagnosis methods for silent data error (SDE) characterization and mitigation on AI and general-purpose compute platforms, including heterogeneous systems and large-scale server clusters.
Role type
System reliability research intern
Builds
Detection/diagnosis methods for data integrity and platform robustness across HW/FW/OS/runtime stack
Domain
Semiconductor hardware reliability, AI infrastructure, system architecture
Deliverable
research
Required skills
Python programming, Linux scripting, data analysis (pandas/numpy/matplotlib/SQL), computer architecture fundamentals, RAS concepts (ECC/CRC/parity/scrubbing/checkpoints), research hypothesis formulation
Preferred skills
GPU/accelerator stack knowledge, distributed training/inference experience, log analytics, Github Copilot usage
Technologies
Python, Linux, pandas, numpy, matplotlib, SQL, ECC, CRC, EDAC, PCIe, CXL
Responsibilities
Collect and analyze platform telemetry/error logs to identify failure patterns; Design and execute fault injection/stress tests to reproduce silent data corruption; Research in-field scan and lockstep mode features for error detection; Integrate Silicon Lifecycle Management (SLM) solutions with telemetry for health monitoring; Develop scripts/tools to automate data processing and experiment orchestration; Evaluate mitigation techniques (retry/recovery/checkpoint/restart) and quantify impact; Collaborate with cross-functional teams to trace error propagation and document findings
Seniority
Intern, research-focused