CareerPlanSign in

Machine Learning Performance Engineer

New York💼 Full-time🗓 2026-07-30 → 2026-09-26

Core

Optimizing the performance of machine learning models for both training and inference in real-time trading systems.

Role type

Senior IC machine learning performance engineer (low-level systems)

Builds

Efficient large-scale training pipelines and low-latency/high-throughput inference systems for trading

Domain

Financial technology / High-performance computing / GPU systems

Deliverable

production ML models

Required skills

Low-level GPU programming (PTX, SASS, warps, Tensor Cores), CUDA optimization, Distributed training algorithms (NCCL, MPI), High-performance networking (Infiniband, RoCE, NVLink), Systems debugging (CUDA GDB, NSight), CUDA libraries (Triton, CUTLASS, cuDNN)

Preferred skills

Experience with storage systems and host-level optimization

Technologies

CUDA, PTX, SASS, NCCL, MPI, Infiniband, RoCE, NVLink, Triton, CUTLASS, CUB, Thrust, cuDNN, cuBLAS, NSight Systems, NSight Compute

Responsibilities

Debug training runs end-to-end, optimize GPU memory hierarchy and cache usage, tune network throughput for GPU clusters, analyze latency vs goodput at the hardware level

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.