CareerPlanGet AI match score →

Member of Technical Staff - ML Performance

New York💼 Full-time🗓 2026-04-21 → 2026-08-01

Core

Building a new infrastructure layer for AI to enable instant GPU access, sub-second container starts, and native storage for low-latency inference and fine-tuning.

Role type

Senior IC ML performance engineer (infrastructure)

Builds

Modal's container runtime and infrastructure layer for serving language and diffusion models

Domain

Cloud infrastructure + Machine Learning

Deliverable

production ML models

Required skills

high-performance code, PyTorch, inference engines (vLLM, TensorRT), Nvidia GPU architecture, CUDA, ML performance engineering, Linux kernel, file systems, containers

Preferred skills

open-source contributions

Technologies

Modal, PyTorch, vLLM, TensorRT, CUDA, Linux

Responsibilities

Contributing to open-source projects, optimizing Modal's container runtime for higher throughput and lower latency, debugging SM occupancy issues, rewriting algorithms to be compute-bound, eliminating host overhead

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗