Senior Systems Engineer – Performance & Reliability
Core
Evaluating large-scale Linux systems from single racks to data-centre scale to determine production readiness and reliability.
Role type
Senior Systems Engineer (Performance & Reliability)
Builds
Execution systems and workloads for measuring system behaviour at scale
Domain
AI compute infrastructure / Semiconductor / Datacenter scale Linux systems
Deliverable
production ML models | infrastructure
Required skills
Linux-based environments, Python, automation and CI/CD systems, experiment design, result interpretation
Preferred skills
large-scale or distributed systems, performance/reliability testing, pytest, system behaviour analysis under load, containerisation and orchestration
Technologies
GitLab CI, Jenkins, GitHub Actions, Docker, Kubernetes, OpenStack, pytest
Responsibilities
Expanding measurement coverage from small clusters to full racks, designing workloads that expose system behaviour, building systems to run experiments across large clusters, interpreting results and defining what 'good enough' looks like
Seniority
Senior, hands-on IC