Member of Technical Staff - Inference Runtime
Core
Build production container runtime infrastructure to accelerate and secure large-scale ML inference and training workloads across a heterogeneous fleet.
Role type
Senior IC systems engineer (container runtime & distributed systems)
Builds
Production container runtime stack, multi-GPU snapshotting pipelines, and low-latency data paths for ML workloads (via careerplan.io/jobs/6195f552-ca6f-4540-b7f6-e7e6054e47ad-member-of-technical-staff-inference-runtime-at-modal)
Domain
Cloud infrastructure, distributed systems, GPU computing, Linux kernel internals
Deliverable
production ML models
Required skills
Linux systems internals, Rust, Go, container runtime development, GPU driver integration, performance profiling, kernel debugging, distributed systems architecture
Preferred skills
gVisor, RDMA, CUDA/ROCm, FUSE/EROFS, checkpoint/restore systems, open-source contributions
Technologies
Rust, Go, Linux kernel, gVisor, RDMA, NVIDIA/AMD GPUs, EROFS, FUSE, CUDA, ROCm
Responsibilities
Optimize container startup and checkpoint/restore for large ML workloads; Design zero-copy data paths between memory, storage, and runtime; Debug complex failures involving kernel behavior, GPU drivers, and container isolation; Extend sandboxed runtime support for new GPU hardware and drivers; Safely roll out runtime changes across a heterogeneous fleet
Seniority
Senior, hands-on IC
