Senior Machine Learning Engineer (Large Systems)
Seniority
Senior, Staff
Domain
AI, Machine Learning, Semiconductor, Software, Hardware
Required skills
Bachelor/Master's/PhD or equivalent experience in Machine Learning, Computer Science, Maths, Data Science, or related field, Proficiency in deep learning frameworks like PyTorch/JAX, Strong Python or C++ software development skills, Expertise in deep learning from model training to optimisation and evaluation, Experience in distributed training or inference of ML models across 64+ accelerators, Capable of designing, executing and reporting from ML experiments, Developed deep understanding of performance bottlenecks and how to overcome them, Ability to move quickly in a dynamic environment, Enjoy cross-functional work collaborating with other teams, Strong communicator - able to explain complex technical concepts to different audiences
Preferred skills
Experience in one or more of: MLOps for Kubernetes-based clusters, Building production systems with large language models, Efficient computing based on low-precision arithmetic, Experience writing C++/Triton/CUDA kernels for performance optimisation of ML models, Familiarity with HPC systems and networking including Infiniband, NVLink, RoCE technologies, Have contributed to open-source projects or published research papers in relevant fields, Knowledge of cloud computing platforms, Keen to present, publish and deliver talks in the AI community
Technologies
PyTorch, JAX, Python, C++, C++, Triton, CUDA, Kubernetes, HPC systems, Infiniband, NVLink, RoCE, cloud computing platforms
Responsibilities
Implement latest machine learning models and optimise them for performance and accuracy, scaling to 1000s of accelerators, Test and evaluate new internal software releases, provide feedback to software engineering teams, make necessary code fixes, and conduct code reviews, Benchmark models and key ML techniques to identify performance bottlenecks and improve model efficiency, Design and conduct experiments on novel AI methods, implement them and evaluate results, Collaborate with Research, Software, and Product teams to define, build, and test Graphcore’s next generation of AI hardware, Engage with AI community and keep in touch with the latest developments in AI
