Senior AI Systems Performance Engineer
Core
Optimizing and scaling state-of-the-art foundation models on SambaNova's reconfigurable dataflow platform to deliver world-record performance for large-scale AI inference.
Role type
Senior IC ML performance engineer (systems/hardware)
Builds
High-performance AI inference solutions on SambaNova Suite
Domain
AI infrastructure / Systems performance / Hardware-software co-design
Deliverable
production ML models
Required skills
Deep learning model development and performance optimization, Compiler/runtime/kernel-level optimization, Software-hardware co-design, Python/C++ proficiency, ML framework expertise (PyTorch/TensorFlow/JAX), Real-world ML pipeline analysis
Preferred skills
LLM/multimodal model training and inference, Large-scale distributed training and high-throughput inference systems, Quantization/graph optimization/kernel fusion, DeepSpeed/Megatron/vLLM/TensorRT experience, GPU programming (CUDA/Triton/OpenCL), Memory hierarchy optimization
Technologies
SambaNova Suite, SN40L chip, DeepSeek R1, GPT OSS, PyTorch, TensorFlow, JAX, DeepSpeed, Megatron, vLLM, TensorRT, CUDA, Triton, OpenCL, cuDNN, cuBLAS
Responsibilities
Bring up and optimize cutting-edge foundation models on the SambaNova platform, Profile and enhance model performance across compiler, runtime, and hardware layers, Collaborate with ML/compiler/runtime/hardware teams to deliver co-designed applications, Integrate advances in model architecture, quantization, scheduling, and memory optimization, Develop robust end-to-end inference solutions, Identify performance bottlenecks and propose dataflow or scheduling optimizations
Seniority
Senior, hands-on IC