Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization
Core
Own the inference and optimization layer for AI in agentic workflows, fine-tuning models and building runtimes to ensure fast, efficient, production-grade performance on real hardware.
Role type
Senior Staff / Principal Machine Learning Scientist (AI Inference & Optimization)
Builds
Model inference path, task execution runtime, and hardware acceleration support for bounded AI tasks.
Domain
Cloud security, agentic AI, large-scale model inference
Deliverable
production ML models
Required skills
Fine-tuning (LoRA/QLoRA), quantization (GGUF/AWQ/GPTQ), inference runtimes (vLLM/SGLang/TensorRT-LLM/ONNX Runtime/llama.cpp/MLX/CoreML), transformer internals (KV cache, attention, batching, memory footprint), Python, C++ interop
Preferred skills
On-device or edge inference experience, agentic coding systems (Claude Code, Pi, Codex)
Technologies
vLLM, SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, MLX, CoreML, GGUF, AWQ, GPTQ, LoRA, QLoRA
Responsibilities
Build and optimize model inference path (quantization, KV-cache, batching, latency/memory/throughput tuning), fine-tune and evaluate models for bounded tasks, design and grow task execution runtime, drive hardware acceleration and sparsity support, partner with systems/backend engineers to ship capabilities end-to-end
Seniority
Senior Staff / Principal, hands-on IC