Staff Software Engineer – ML Inference Runtime
Core
Develop C++, C, and Python software integrating machine-learning frameworks and LLM runtimes with Arm hardware acceleration technologies.
Role type
Staff Software Engineer (ML Inference Runtime)
Builds
Integrations between ML frameworks, runtimes, and hardware-acceleration technologies for Arm platforms
Domain
Hardware acceleration, Machine Learning, LLM inference
Deliverable
production ML models
Required skills
C++, C, Python, LLM inference (tokenization, attention, KV caches, batching, quantisation), Git, CI/CD, automated testing, technical leadership
Preferred skills
llama.cpp, workload profiling, compiler/ML graph technologies, Vulkan/GPU compute APIs, AI-assisted development tools
Technologies
Arm, llama.cpp, C++, C, Python, Git, GitLab, Vulkan, GPU compute
Responsibilities
Design and implement integrations between ML frameworks, runtimes, and hardware-acceleration technologies; analyse and improve performance of ML and LLM workloads; lead significant work from design through delivery; mentor junior engineers
Seniority
Staff, hands-on IC with leadership (via careerplan.io/jobs/100035000624-staff-software-engineer-ml-inference-runtime-at-arm)
