Runtime Engineer
Core
Build the host-side interface library and runtime stack for custom silicon designed for large-language-model inference and training.
Role type
Senior IC Runtime Engineer (LLM inference/training stack)
Builds
Host-side interface library, executable format contracts, custom-kernel ABI, Python bindings, LLM inference serving stack, and cluster orchestration primitives.
Domain
Hardware/Software co-design for LLM inference and training
Deliverable
production ML models | infrastructure
Required skills
Systems programming (Rust/C/C++/Go), Python interop (PyO3/ctypes/pybind11), ABI/contract design, accelerator programming models (CUDA/ROCm/oneAPI/TPU), ML-systems literacy
Preferred skills
LLM inference internals (vLLM/TensorRT-LLM/SGLang), Rust depth (proc macros/unsafe), custom allocator design, ML framework integration, profiler/tracing infrastructure
Technologies
PyO3, CUDA, ROCm, oneAPI Level Zero, TPU, DLPack, numpy, vLLM, TensorRT-LLM, SGLang, PyTorch, JAX/XLA, ONNX runtime, Perfetto, Nsight
Responsibilities
Build host-side interface library for device memory management and sync primitives; Own and extend executable format and compiler-runtime contracts; Design custom-kernel ABI and host-side marshaling layers; Build LLM inference serving stack with paged KV cache and request scheduling; Bring up interconnect topology and failure-detection paths; Design chip exposure for profilers and debuggers.
Seniority
Senior, hands-on IC
