System Engineer (Token Factory)
Core
Building low-level kernels and runtime components for AI inference on massive GPU clusters to optimize performance for foundation models.
Role type
Senior IC systems engineer (GPU inference)
Builds
High-performance inference engines and runtime components for Nebius Cloud's GPU platform.
Domain
Cloud infrastructure + AI inference optimization
Deliverable
production ML models
Required skills
C++, GPU programming, low-level high-performance coding, memory management, systems-level software development, profiling and debugging, hardware architecture integration
Preferred skills
Experience with Hopper, Blackwell, or Rubin GPU architectures
Technologies
C++, GPU platforms (Hopper, Blackwell, Rubin)
Responsibilities
Develop and optimize low-level kernels and runtime components for AI inference; Improve performance of inference engines on GPU platforms; Profile and debug system-level and hardware-level performance issues; Integrate support for new hardware architectures; Collaborate with ML and backend teams to optimize end-to-end execution
Seniority
Senior, hands-on IC