Intern - ML Inference Performance Engineer
Core
Develop repeatable evaluation methodologies and tooling to benchmark ML inference performance across hardware and software platforms, shaping product decisions.
Role type
Intern ML Inference Performance Engineer
Builds
Benchmarking tooling, performance dashboards, and evaluation reports for AI accelerator platforms
Domain
AI Infrastructure / ML Inference Optimization
Deliverable
production ML models
Required skills
Python development, C/C++ programming, end-to-end computer vision pipeline experience, benchmarking concepts, inference tools/APIs/SDKs (e.g., TensorRT), deep learning model concepts (quantization, ONNX, PyTorch), Linux/Bash/Docker proficiency, embedded host experience, Git version control
Preferred skills
GStreamer knowledge, basic GUI design experience
Technologies
TensorRT, ONNX, PyTorch, GStreamer, Linux, Docker
Responsibilities
Develop and improve internal benchmarking tools for throughput, latency, and power; research and evaluate AI accelerator products from various vendors; characterize full inference pipelines and host-device transaction overhead; set up and maintain lab hosts across multiple hardware platforms; synthesize findings into structured reports for engineering decisions
Seniority
Intern