Senior Systems Engineer – Performance & Reliability (Analysis)
Core
Turn large-scale system measurements into decisions by analyzing performance data from clusters of machines to determine if systems are stable enough for production.
Role type
Senior Systems Engineer (Performance & Reliability Analysis)
Builds
Production-ready AI compute systems and infrastructure
Domain
Semiconductor hardware, AI compute stack, distributed systems
Deliverable
production ML models | infrastructure
Required skills
Linux-based environment expertise, Python programming, automation and CI/CD systems, experimental design, data interpretation, engineering judgement
Preferred skills
Large-scale or distributed systems experience, performance/reliability testing, pytest, system behavior analysis under load, containerization and orchestration, C++
Technologies
GitLab CI, Jenkins, GitHub Actions, Docker, Kubernetes, OpenStack, pytest
Responsibilities
Run workloads across clusters and collect detailed performance data, interpret results to assess system stability and production readiness, design and execute experiments to produce meaningful results, communicate findings clearly to support engineering decisions
Seniority
Senior, hands-on IC