Senior and/or Principal Software Engineer - Perfomance
Core
Identify and drive improvements to end-to-end inference performance of OpenAI and other state-of-the-art LLMs; build SW tooling to enable insights into performance opportunities ranging from the model level to the systems and silicon level.
Role type
Senior/Principal Software Engineer (LLM Performance & Optimization)
Builds
SW tooling for LLM inference performance analysis, model porting on new GPUs, and AI/DNN/LLM frameworks.
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
C, C++, C#, Java, JavaScript, Python, GPU architecture, HW neural net acceleration, end-to-end performance analysis, GPU profiling, DNN/LLM inference, PyTorch, Tensorflow, ONNX Runtime, CUDA, ROCm, Triton
Preferred skills
None stated
Technologies
Nvidia GPUs, AMD GPUs, PyTorch, Tensorflow, ONNX Runtime, CUDA, ROCm, Triton
Responsibilities
Optimize and monitor performance of LLMs; build SW tooling to enable insights into performance opportunities; enable fast time to market of LLMs/models and their deployments at scale; design, implement, and test functions or components for AI/DNN/LLM frameworks; speed up/reduce complexity of key components/pipelines to improve performance and/or efficiency of systems.
Seniority
Senior/Principal, hands-on IC