Machine Learning Engineer
Core
Develop evaluation frameworks, benchmarks, and metrics to measure and improve the performance of LLM and agent-based AI systems for the mining industry.
Role type
Machine Learning Engineer (LLM evaluation & benchmarking)
Builds
Evaluation frameworks, benchmarks, and test datasets for LLM and agentic AI systems
Domain
Mining industry + Large Language Models (LLM) and Agentic AI
Deliverable
production ML models
Required skills
PyTorch, Python, NumPy, Pandas, Scikit-learn, neural networks, computer vision, natural language processing
Preferred skills
ML evaluation, benchmarking for LLM or agentic systems, deploying ML models to production
Technologies
PyTorch, NumPy, Pandas, Scikit-learn
Responsibilities
Design and implement evaluation frameworks for LLM and agentic AI systems; Develop benchmarks, metrics, and test datasets to assess model quality; Analyse model performance and identify improvement opportunities; Partner with product and engineering teams to define AI success measures
Seniority
Mid-level, hands-on IC
