Software Engineer, Inference - Multi Modal
Core
Building reliable, high-performance infrastructure to serve real-time audio, image, and multimodal AI models at scale.
Role type
Senior IC software engineer (multimodal inference infrastructure)
Builds
Production systems for serving large-scale multimodal models (image, audio, text) with high throughput and low latency
Domain
Artificial Intelligence / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
scaling inference systems for LLMs or multimodal models, GPU-based ML workload optimization, distributed compute and networking, high-throughput data handling, system-level improvements (GPU utilization, tensor parallelism, hardware abstraction), familiarity with inference tooling (vLLM, TensorRT-LLM, custom model parallel systems)
Preferred skills
experience with image generation or audio synthesis models in production, distributed ML training, system-efficient model design
Technologies
vLLM, TensorRT-LLM, GPU clusters, distributed compute frameworks
Responsibilities
Design and implement inference infrastructure for large-scale multimodal models, Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs, Enable experimental research workflows to transition into reliable production services, Contribute to system-level improvements including GPU utilization and hardware abstraction layers
Seniority
Senior, hands-on IC