CareerPlanSign in

Deep Learning Performance Software Intern - 2027

China, Shanghai💼 Internship🗓 2026-09-30 → 2026-10-01

Core

Developing GPU-accelerated deep learning software and highly optimized deep learning kernels using tile-based GPU programming models.

Role type

Deep Learning Performance Software Engineering Intern

Builds

SKILL, Wiki, agent harness, TileGym, Triton CUDA TileIR backend, CUDA Tile

Domain

Deep Learning, GPU Computing, High-Performance Computing

Deliverable

production ML models

Required skills

C/C++ programming, software design, agentic systems (LLM APIs, prompting, tool use, agent loops, planning, reasoning, memory, RAG, skills, MCP, sub-agents, multi-agent architectures, context engineering, harness engineering), performance modeling, profiling, debugging, code optimization, architectural knowledge of CPU and GPU

Preferred skills

Python, MLIR, GPU programming (CUDA or OpenCL), masters or doctoral degree

Technologies

CUDA, Triton, MLIR, C/C++, Python, LLM APIs, RAG, MCP

Responsibilities

Creating and maintaining SKILL, Wiki, and agent harness; Developing TileGym, Triton CUDA TileIR backend, and CUDA Tile; Developing highly optimized deep learning kernels through tile-based GPU programming model; Performing end-to-end performance optimization through tile-based GPU programming model; Conducting performance optimization, analysis, and tuning

Sourced via workday · Listed on CareerPlan, which tracks 912,000+ jobs from 20+ sources.