Machine Learning Engineer
Core
Optimizing large language model inference on HPC clusters and GPU infrastructure.
Role type
Senior Machine Learning Engineer (Inference Optimization)
Builds
Production inference pipelines for large language models
Domain
Cloud Infrastructure / Large Language Models
Deliverable
production ML models
Required skills
Quantization, PEFT, DeepSpeed, ONNX, TensorRT, PyTorch, multi-LoRa, LoRA Exchange, TitanML, HPC cluster management, LLM scaling, GPU data movement
Preferred skills
Experience with OpenAI, Mistral, Claude, LLaMA models
Technologies
PyTorch, ONNX, TensorRT, TitanML, LoRA Exchange
Responsibilities
Implement quantization and optimization techniques for LLMs, manage inference servers, scale GPU usage and data movement, work with multi-node HPC clusters
Seniority
Senior, hands-on IC
Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.