AI Inference Platform Engineer
Core
Build, operate, and optimize the systems that serve large language, vision, multimodal, and embedding models across DRW, providing the firmwide interface to modern AI models from evaluation through production use.
Role type
Senior IC AI Inference Platform Engineer
Builds
Inference runtimes, distributed systems, and production platform for LLMs and multimodal models
Domain
Financial Trading / AI Infrastructure
Deliverable
production ML models
Required skills
LLM inference optimization, NVIDIA GPU architecture expertise, inference runtime development, distributed system design, multi-tenant scheduling, performance profiling, model quality equivalence testing, CI/CD for model serving
Preferred skills
Experience with Hopper/Blackwell GPUs, knowledge of speculative decoding and quantization, Linux systems performance fundamentals
Technologies
TensorRT-LLM, vLLM, SGLang, CUDA, Nsight, DCGM, OpenTelemetry, Prometheus, Grafana
Responsibilities
Optimize LLM inference performance across modern NVIDIA GPU architectures; Build end-to-end performance profiling and observability; Design and optimize KV cache and distributed inference architectures; Own day-0 model onboarding and serving configuration; Measure and monitor quality equivalence across serving configurations; Manage the production serving lifecycle of models; Partner with SRE and platform teams to automate deployment and operation; Optimize model placement and resource allocation; Design and operate multi-tenant scheduling and isolation
Seniority
Senior, hands-on IC