Manager, Machine Learning Engineering
Core
Lead the team responsible for the infrastructure, pipelines, and operational rigor that keep GoFundMe's machine learning and AI systems reliable, scalable, and safe in production.
Role type
Manager, Machine Learning Engineering (MLOps and AI Operations)
Builds
Production ML/AI infrastructure, pipelines, feature stores, model serving, and monitoring/observability systems
Domain
Non-profit / Community fundraising / Generative AI and Traditional ML
Deliverable
production ML models | infrastructure
Required skills
Python, PyTorch, TensorFlow, Scikit-learn, AWS, Databricks, Docker, Kubernetes, Terraform, Snowflake, SQL, Spark, CI/CD, real-time model serving, ML monitoring, team leadership, hiring, architecture design
Preferred skills
Generative AI/LLM infrastructure experience, advanced degree in CS/Statistics/Data Science
Technologies
Python, AWS, Databricks, Docker, Kubernetes, Terraform, Snowflake, GitHub, PyTorch, TensorFlow, Scikit-learn
Responsibilities
Own the reliability, scalability, and operational health of ML/AI production systems; Lead, hire, and grow a team of ML/AI operations engineers; Partner with data science and ML engineering teams to streamline model deployment; Establish ML operational excellence org-wide; Build and mature on-call processes, SLOs/SLAs, and postmortem practices; Drive operational strategy for generative AI systems; Manage vendor and platform relationships; Report on team health, system reliability metrics, and operational risk
Seniority
Manager, 1-3+ years of management experience
