Machine Learning Performance Engineer
Core
Optimizing the performance of machine learning models for both training and inference in real-time trading systems.
Role type
Senior IC machine learning performance engineer (low-level systems)
Builds
Efficient large-scale training pipelines and low-latency/high-throughput inference systems for trading
Domain
Financial technology / High-performance computing / GPU systems
Deliverable
production ML models
Required skills
Low-level GPU programming (PTX, SASS, warps, Tensor Cores), CUDA optimization, Distributed training algorithms (NCCL, MPI), High-performance networking (Infiniband, RoCE, NVLink), Systems debugging (CUDA GDB, NSight), CUDA libraries (Triton, CUTLASS, cuDNN)
Preferred skills
Experience with storage systems and host-level optimization
Technologies
CUDA, PTX, SASS, NCCL, MPI, Infiniband, RoCE, NVLink, Triton, CUTLASS, CUB, Thrust, cuDNN, cuBLAS, NSight Systems, NSight Compute
Responsibilities
Debug training runs end-to-end, optimize GPU memory hierarchy and cache usage, tune network throughput for GPU clusters, analyze latency vs goodput at the hardware level
Seniority
Senior, hands-on IC