Senior ML Engineer (Token Factory)
Core
Building low-level kernels and runtime components for AI inference on massive GPU clusters to optimize performance for foundation models.
Role type
Senior IC systems/ML engineer (GPU inference)
Builds
High-performance inference engines and runtime components for text, vision, and multimodal foundation models.
Domain
Cloud infrastructure / GPU computing / AI systems
Deliverable
production ML models
Required skills
C++ or GPU programming, systems-level software development, profiling and debugging, memory management, hardware architecture integration
Preferred skills
Experience with Hopper/Blackwell/Rubin GPU architectures, end-to-end execution optimization
Technologies
C++, GPU platforms, Hopper, Blackwell, Rubin
Responsibilities
Develop and optimize low-level kernels and runtime components for AI inference; Improve performance of inference engines on GPU platforms; Profile and debug system-level and hardware-level performance issues; Integrate support for new hardware architectures; Collaborate with ML and backend teams to optimize end-to-end execution