Staff Python / PyTorch Developer — Frontend Inference Compiler – Dubai
Core
Design and implement APIs, machine learning features, and tools enabling state-of-the-art generative AI models to run efficiently on custom hardware.
Role type
Senior Software Engineer (Inference ML Engineering)
Builds
Scalable, high-performance inference solutions for generative AI models
Domain
AI hardware, Large Language Models (LLMs), Multimodal AI
Deliverable
production ML models
Required skills
Python, C++, PyTorch, LLM serving frameworks (vLLM, SGLang, TensorRT-LLM), software architectural patterns, performance optimization, observability, automated testing
Preferred skills
Experience with structured outputs, biased sampling, predicted outputs, multimodal inference (image, audio, video), agile development practices
Technologies
Python, C++, PyTorch, vLLM, SGLang, TensorRT-LLM
Responsibilities
Design and implement ML features to improve generative AI model performance; Design and implement high-throughput, low-latency multimodal inference models; Maintain scalable serving backend; Scale inference service via observability; Optimize software for high throughput and low latency; Analyze and improve latency, throughput, memory usage, and compute efficiency; Build and maintain robust automated test suites; Lead cross-functional initiatives
Seniority
Senior, hands-on IC with technical guidance