Member of Technical Staff - Inference Systems
Core
Designing and building the engine layer for running AI models in production, including benchmarking infrastructure for performance and quality evaluation.
Role type
Senior IC inference systems engineer
Builds
Inference engine layer, benchmark suites, and partner verification pipelines
Domain
AI infrastructure, model deployment, performance engineering
Deliverable
production ML models
Required skills
C++, Python, inference frameworks (llama.cpp, ONNX, MLX), benchmark design, model porting, quantization, memory layout optimization
Preferred skills
Edge inference constraints, external partner technical validation, numerical correctness verification
Technologies
llama.cpp, ONNX Runtime, MLX
Responsibilities
Design and build benchmark suites for inference performance and model quality; Run external partner verifications against benchmarks; Port models onto different runtimes and frameworks; Maintain and extend the inference engine layer; Make benchmark results explainable and verifiable
Seniority
Senior, hands-on IC