CareerPlanGet AI match score →

Cloud and AI System Intern

PRC, Shanghai💼 Internship🗓 2026-04-24 → 2026-07-30

Core

Research and prototype detection/diagnosis methods for silent data error (SDE) characterization and mitigation on AI and general-purpose compute platforms, including heterogeneous systems and large-scale server clusters.

Role type

System reliability research intern

Builds

Detection/diagnosis methods for data integrity and platform robustness across HW/FW/OS/runtime stack

Domain

Semiconductor hardware reliability, AI infrastructure, system architecture

Deliverable

research

Required skills

Python programming, Linux scripting, data analysis (pandas/numpy/matplotlib/SQL), computer architecture fundamentals, RAS concepts (ECC/CRC/parity/scrubbing/checkpoints), research hypothesis formulation

Preferred skills

GPU/accelerator stack knowledge, distributed training/inference experience, log analytics, Github Copilot usage

Technologies

Python, Linux, pandas, numpy, matplotlib, SQL, ECC, CRC, EDAC, PCIe, CXL

Responsibilities

Collect and analyze platform telemetry/error logs to identify failure patterns; Design and execute fault injection/stress tests to reproduce silent data corruption; Research in-field scan and lockstep mode features for error detection; Integrate Silicon Lifecycle Management (SLM) solutions with telemetry for health monitoring; Develop scripts/tools to automate data processing and experiment orchestration; Evaluate mitigation techniques (retry/recovery/checkpoint/restart) and quantify impact; Collaborate with cross-functional teams to trace error propagation and document findings

Seniority

Intern, research-focused

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗