Member of Technical Staff - Inference
Core
Develop training & inference benchmarks and system modeling for AI hardware and clusters.
Role type
Individual contributor ML infrastructure engineer (benchmarking & system modeling)
Builds
Training & inference benchmarks, system models for AI compute clusters
Domain
Semiconductor supply chain, AI hardware, datacenter infrastructure
Deliverable
production ML models | dashboards & analysis
Required skills
ML frameworks (PyTorch, JAX), transformer architecture, LLM parallelism, CUDA parallel programming, Python, NCCL, technical research, system modeling
Preferred skills
Experience with NVIDIA H100, AMD Mi300X, Google TPUs, AWS Trainium2, distributed system scaling
Technologies
PyTorch, JAX, vLLM, SGLang, NCCL, Python
Responsibilities
Conduct training & inference performance benchmarks across various AI hardware; Author detailed technical research reports analyzing benchmark results; Develop comprehensive system modelling using Python & NCCL for existing & future AI compute clusters; Establish and maintain strategic partnerships with neocloud providers & AI chip manufacturers; Stay current on emerging trends by attending major industry & academic conferences
Seniority
Individual contributor, all experience levels