Senior Inference Engineer - AI
Core
Productionizing, optimizing, and scaling AI and LLM workloads to power Thomson Reuters' AI-driven products across a multi-cloud footprint.
Role type
Senior Inference Engineer (AI Infrastructure)
Builds
Production-grade inference pipelines, containerized workloads, and optimized model routing for enterprise AI services.
Domain
Enterprise AI, Cloud Infrastructure, Machine Learning Operations
Deliverable
production ML models
Required skills
LLM inference optimization, GPU programming (CUDA), inference runtimes (TensorRT, ONNX Runtime), deep learning frameworks (PyTorch, TensorFlow), Python, systems programming (C++)
Preferred skills
Quantization, pruning, distillation, heterogeneous hardware tuning, vector search systems (OpenSearch)
Technologies
AWS, Azure, GCP, OCI, Kubernetes, Snowflake, OpenSearch, PyTorch, TensorFlow, CUDA, TensorRT, ONNX Runtime
Responsibilities
Optimize LLMs and ML models for high-performance inference using quantization and hardware tuning; Deploy and scale inference workloads on GPUs across multi-cloud environments; Implement routing and failover strategies for model traffic; Integrate models into production APIs; Profile performance and eliminate bottlenecks; Build containerized inference pipelines using Kubernetes; Ensure compliance with AI standards for deployment and monitoring.
Seniority
Senior, hands-on IC