CareerPlanSign in

Machine Learning Engineer, Reliability

💼 Full-time🗓 2026-06-24 → 2026-09-25

Core

Own the reliability, security, and safety of a large fleet of generative media model APIs serving production traffic at scale.

Role type

Senior hybrid Machine Learning Engineer / Site Reliability Engineer

Builds

High-performance inference, orchestration, and observability for image, video, and audio model APIs

Domain

Generative AI / Infrastructure / Cloud Systems

Deliverable

production ML models

Required skills

distributed systems, networking, observability, incident management, canary releases, shadow testing, automated rollbacks, validation gates, capacity planning, autoscaling, GPU fleet efficiency, abuse detection, content moderation, safety classifiers

Preferred skills

diffusion models, transformers, security practices for ML systems, trust & safety engineering

Technologies

Python, torch, diffusers, Kubernetes, fal Python SDK

Responsibilities

Own availability, latency, and throughput SLOs; build monitoring and alerting for ML-specific failures; harden model deployment workflows; drive security posture and abuse prevention; operationalize safety systems and content moderation; lead incident response and postmortems; improve capacity planning and GPU efficiency; partner with model teams on reliability requirements

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.