Manager of Machine Learning - AI Modeling and Operation
Core
Lead the team responsible for ML infrastructure, model operations, and AI quality to ensure every AI feature at Workiva is reliable, observable, and deployable.
Role type
Engineering Manager for ML Infrastructure and MLOps
Builds
Operational backbone for enterprise AI, including model lifecycle management, evaluation frameworks, observability, and model routing across frontier providers
Domain
Enterprise Finance & Risk / Generative AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
ML pipeline orchestration, Kubernetes, microservices, infrastructure-as-code, production ML system reliability, team leadership, incident response, cloud operations (AWS/Azure/GCP)
Preferred skills
Generative AI concepts (RAG, Agentic frameworks), model evaluation systems, GPU/model serving cost optimization, observability tooling (Datadog, Prometheus, Grafana)
Technologies
ClearML, Kubeflow, Airflow, Kubernetes, AWS Bedrock, Azure OpenAI, Google Cloud, Datadog, Prometheus, Grafana
Responsibilities
Own the ML model lifecycle from training to deployment and monitoring; Build and maintain CI/CD for ML with automated testing and evaluation; Drive observability across AI services including latency tracking and drift detection; Lead incident response for ML-related production issues; Lead and grow a team of machine learning engineers; Partner with product and engineering squads to ensure AI services are production-ready
Seniority
Manager, hands-on technical leadership
