Member of Technical Staff - ML Performance
Core
Building a new infrastructure layer for AI to enable instant GPU access, sub-second container starts, and native storage for low-latency inference and fine-tuning.
Role type
Senior IC ML performance engineer (infrastructure)
Builds
Modal's container runtime and infrastructure layer for serving language and diffusion models
Domain
Cloud infrastructure + Machine Learning
Deliverable
production ML models
Required skills
high-performance code, PyTorch, inference engines (vLLM, TensorRT), Nvidia GPU architecture, CUDA, ML performance engineering, Linux kernel, file systems, containers
Preferred skills
open-source contributions
Technologies
Modal, PyTorch, vLLM, TensorRT, CUDA, Linux
Responsibilities
Contributing to open-source projects, optimizing Modal's container runtime for higher throughput and lower latency, debugging SM occupancy issues, rewriting algorithms to be compute-bound, eliminating host overhead
Seniority
Senior, hands-on IC