Machine Learning Engineer, Reliability
Core
Own the reliability, security, and safety of a large fleet of generative media model APIs serving production traffic at scale.
Role type
Senior hybrid Machine Learning Engineer / Site Reliability Engineer
Builds
High-performance inference, orchestration, and observability for image, video, and audio model APIs
Domain
Generative AI / Infrastructure / Cloud Systems
Deliverable
production ML models
Required skills
distributed systems, networking, observability, incident management, canary releases, shadow testing, automated rollbacks, validation gates, capacity planning, autoscaling, GPU fleet efficiency, abuse detection, content moderation, safety classifiers
Preferred skills
diffusion models, transformers, security practices for ML systems, trust & safety engineering
Technologies
Python, torch, diffusers, Kubernetes, fal Python SDK
Responsibilities
Own availability, latency, and throughput SLOs; build monitoring and alerting for ML-specific failures; harden model deployment workflows; drive security posture and abuse prevention; operationalize safety systems and content moderation; lead incident response and postmortems; improve capacity planning and GPU efficiency; partner with model teams on reliability requirements
Seniority
Senior, hands-on IC