Senior Systems Engineer – Performance & Reliability
Core
Designing and running measurements on large-scale Linux clusters to determine system performance, reliability, and production readiness for AI compute stacks.
Role type
Senior Systems Engineer (Performance & Reliability)
Builds
Measurement infrastructure and analysis pipelines for datacentre-scale systems
Domain
AI Compute / Semiconductor / Datacenter Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, Python, distributed systems, performance analysis, CI/CD automation, experiment design, result interpretation
Preferred skills
High-performance computing, pytest, GitLab CI/Jenkins/GitHub Actions
Technologies
Linux, Python, pytest, GitLab CI, Jenkins, GitHub Actions
Responsibilities
Running measurements on large-scale Linux clusters (rack-level and beyond), Using and extending tools such as pytest for measurement execution, Measuring compute, network, and ML workload performance, Analysing variability and repeatability of results
Seniority
Senior, hands-on IC