Performance Engineer
Core
Define and measure scale-up fabric performance for rack-scale AI infrastructure by building roofline models, benchmarks, and end-to-end workload studies for GPU clusters.
Role type
Senior Performance Engineer (AI Infrastructure)
Builds
Roofline models, performance benchmarks, and scalability studies for Astera Labs' Scorpio scale-up fabric switches.
Domain
AI Infrastructure / High-Performance Computing / Datacenter Networking
Deliverable
production ML models | product features | dashboards & analysis
Required skills
GPU cluster benchmarking, roofline modeling, system-level debugging, parallel algorithms, datacenter networking (PCIe, Ethernet), Python scripting, automated test infrastructure
Preferred skills
Scale-up fabric expertise (UALink, PCIe Gen 6/7), LLM/MoE workload patterns, competitive performance analysis, executive technical communication
Technologies
NVBandwidth, NCCL, CUDA, MPI, Confluence, Python
Responsibilities
Establish theoretical and measured roofline models for scale-up fabric; build and maintain baseline performance benchmarks using industry-standard tools; quantify impact of fabric features against baselines using synthetic and real inference workloads; run end-to-end inference workloads to evaluate fabric scalability; design and maintain automated lab infrastructure and test pipelines; partner with architecture, firmware, and marketing teams to influence design decisions and support customer engagements.
Seniority
Senior, hands-on IC