Staff Machine Learning Ops Engineer
Core
Define and evolve the ML platform enabling machine learning at company scale for personalized learning, marketplace intelligence, and GenAI products.
Role type
Staff Machine Learning Ops Engineer (Platform & Infrastructure)
Builds
Scalable, secure, observable ML platform and GenAI/LLM infrastructure serving millions of learners and tutors globally.
Domain
EdTech, Machine Learning Operations, Generative AI
Deliverable
production ML models | infrastructure
Required skills
Large-scale ML platform architecture, cloud-native infrastructure, distributed training and inference, CI/CD for ML, observability, LLM/GenAI infrastructure, mentoring senior engineers
Preferred skills
GCP or AWS, Kubernetes, infrastructure-as-code, LangChain, LlamaIndex, vector stores, prompt evaluation
Technologies
GCP, AWS, Kubernetes, LangChain, LlamaIndex, vector stores
Responsibilities
Define technical vision and roadmap for the ML platform; Lead architecture across the full ML lifecycle; Design cloud-native infrastructure for distributed training and inference; Set technical direction for ML CI/CD pipelines; Establish observability standards for ML systems; Lead evolution of GenAI and LLM platform capabilities; Mentor engineers and raise technical standards
Seniority
Staff, hands-on IC with significant leadership and mentorship