Machine Learning Engineer
Core
Develop SOTA deep learning software and model optimization algorithms for LLM inference, focusing on compression, quantization, and speculative decoding to accelerate AI for enterprises.
Role type
Senior IC machine learning engineer (LLM inference optimization)
Builds
Production-ready LLM inference systems, model compression pipelines, and speculative decoding frameworks for enterprise deployment
Domain
Artificial Intelligence / Large Language Models / Model Optimization
Deliverable
production ML models
Required skills
Deep learning fundamentals, LLM inference optimization, tensor math libraries (PyTorch, NumPy), Python, algorithm design, linear algebra, mathematical modeling
Preferred skills
Experience with model quantization and pruning, speculative decoding frameworks, hardware profiling (CPU/GPU), open-source contribution
Technologies
PyTorch, NumPy, vLLM, LLM-compressor
Responsibilities
Design and implement model compression pipelines using quantization and pruning; Develop and maintain speculative decoding frameworks to improve inference speed; Profile and optimize end-to-end LLM performance including memory, latency, and throughput; Collaborate with research scientists to translate experimental ideas into production systems; Contribute to open-source projects and code reviews; Mentor team members
Seniority
Senior, hands-on IC