Full Stack ML Efficiency & Observability
Core
Design and develop features for a capacity management portal and provide visibility into model performance and quality across an ML fleet.
Role type
Full Stack ML Efficiency & Observability Engineer
Builds
Internal tooling, capacity management portal, and observability interfaces for ML training and inference fleets.
Domain
Machine Learning Operations (MLOps), Generative AI infrastructure, Web Development
Deliverable
product features
Required skills
JavaScript, TypeScript, React, HTML, CSS, browser internals, web performance, accessibility, cross-browser compatibility, C/C++/C#/Java/Python, Generative AI tools, capacity management, efficiency management, ML training/inference, software architecture, technical project leadership
Preferred skills
Experience with Visual Studio/VS Code, translating functional requirements into intuitive interfaces, implementing best software development practices
Technologies
React, TypeScript, JavaScript, HTML, CSS, Visual Studio, Visual Studio Code
Responsibilities
Design and develop features for the capacity management portal; Build visibility into model performance and quality across the ML fleet; Integrate with backend APIs from schedulers to training frameworks; Contribute to the development of internal tooling and infrastructure; Ensure code quality and embody company culture.
Seniority
Mid-Senior, hands-on IC