CareerPlanSign in

Member of Technical Staff - Frontier System Modelling

San Francisco, CA, New York City, NY🌐 Remote💼 Full-time🗓 2026-09-20 → 2026-09-26

Core

Develop complex system models in Python to predict performance (MFU, tok/s/gpu) on 100k+ chip AI clusters for frontier LLM training and inference.

Role type

Senior IC system modelling engineer (AI infrastructure)

Builds

Performance prediction models, roofline curves, and cost-of-ownership estimates for next-gen AI chips

Domain

Semiconductor supply chain + AI infrastructure

Deliverable

production ML models

Required skills

Python, parallelism strategies (prefill, expert/tensor/pipeline/sequence), frontier MoE workloads, GEMM operator analysis, Amdahl principles, arithmetic intensity, scaling techniques

Preferred skills

ML engineering, kernel programming, hyperscale cluster modelling, open source contributions

Technologies

Python, AMD, NVIDIA, TPU, Trainium

Responsibilities

Develop Python models to predict AI chip performance; Implement parallelism strategies to generate roofline curves; Build microbenchmarks across vendors to calibrate system models; Build BoM estimates for total cost of ownership and performance per watt

Seniority

Senior, hands-on IC

Sourced via dover · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.