Member of Technical Staff - Frontier System Modelling
Core
Develop complex system models in Python to predict performance (MFU, tok/s/gpu) on 100k+ chip AI clusters for frontier LLM training and inference.
Role type
Senior IC system modelling engineer (AI infrastructure)
Builds
Performance prediction models, roofline curves, and cost-of-ownership estimates for next-gen AI chips
Domain
Semiconductor supply chain + AI infrastructure
Deliverable
production ML models
Required skills
Python, parallelism strategies (prefill, expert/tensor/pipeline/sequence), frontier MoE workloads, GEMM operator analysis, Amdahl principles, arithmetic intensity, scaling techniques
Preferred skills
ML engineering, kernel programming, hyperscale cluster modelling, open source contributions
Technologies
Python, AMD, NVIDIA, TPU, Trainium
Responsibilities
Develop Python models to predict AI chip performance; Implement parallelism strategies to generate roofline curves; Build microbenchmarks across vendors to calibrate system models; Build BoM estimates for total cost of ownership and performance per watt
Seniority
Senior, hands-on IC