AI Systems Performance Engineer
Core
Optimizing Ethernet fabric performance for AI inference and training workloads to ensure maximum throughput and minimum latency.
Role type
Senior AI Systems Performance Engineer
Builds
Performance benchmarks and reports for AI deployment and marketing
Domain
AI/ML infrastructure, Network Engineering, Linux Systems
Deliverable
production ML models | infrastructure
Required skills
Linux system-level tuning, Python, C++, Ethernet fabric optimization, MLPerf benchmarking, distributed system debugging, performance profiling
Preferred skills
RDMA/RoCEv2, CI/CD pipeline automation, Docker, Kubernetes
Technologies
MLPerf, NCCL, Ethernet switches, Linux, Python, C++, PyTorch
Responsibilities
Execute industry-standard AI performance benchmarks (MLPerf, NCCL), tune network parameters for distributed AI clusters, isolate and troubleshoot bottlenecks across OS/hardware/switches, develop automation tools for continuous benchmarking, document methodologies and communicate findings to stakeholders
Seniority
Senior, hands-on IC