CareerPlanGet AI match score →

Generative Ml Engineer

💼 Full-time💰 $24,000–$48,000🗓 2026-07-25

Core

Building scalable ML infrastructure to automate high-end commercial generative media production (image, video, voice) using licensed human likenesses.

Role type

Generative ML Engineer

Builds

Scalable inference pipelines and orchestration systems for licensed generative media

Domain

Generative AI, Creative Technology, Media Production

Deliverable

production ML models

Required skills

Python, Diffusion models, ComfyUI, Serverless GPU, LoRA fine-tuning

Preferred skills

Video/film production interest, Voice cloning, Multi-model orchestration, Azure

Technologies

Modal, ComfyUI, Wan, FLUX, Qwen, Chatterbox, xttsv2, VibeVoice, Fal, Kling, Veo, Runway, ElevenLabs, FastAPI, Azure

Responsibilities

Integrate external generative media APIs; deploy and optimize local models; build inference pipelines; write FastAPI services and inference scripts

Seniority

Mid-level, hands-on IC

Rewrite
## About the role Excury is building the infrastructure for licensed generative media: a platform where brands, producers, and creators can access real people's licensed image, voice, performance, and likeness rights, then generate content safely and at scale. We currently run an advanced AI production studio where our creative team produces high-end commercial work by hand. Your job will be to turn those manual workflows into scalable systems: so others can create high-quality, licensed content with real human characters through Excury. Think what ElevenLabs did for voice. That is the kind of scope we are building across image, video, and voice. The role You are an engineer who follows the generative media space like a beat. You understand the tradeoffs between Wan, Kling, Veo, Runway, FLUX, and the next model drop not because you only read about them, but because you test them, break them, and ship with them. You can go from a Hugging Face release, a model announcement, or a new API to a working internal prototype fast. You find the pace of this space energizing, not exhausting. You will own key parts of our ML infrastructure: integrating external APIs such as Fal, Kling, Veo, Runway, and ElevenLabs; deploying and optimizing local models through ComfyUI and Modal; and building inference pipelines that turn our directors' manual creative workflows into repeatable, scalable systems. You will work with a small team, use AI coding tools heavily, and see your work in production quickly. You will have real ownership, direct impact, and equity upside. ## What you'll work with - Modal for serverless GPU inference and autoscaling - ComfyUI for diffusion pipeline orchestration and programmatic workflow templating - Wan, FLUX, Qwen Image Edit, and the next generation of image/video models - Open source TTS models: Chatterbox, xttsv2, VibeVoice - LoRA persona fine-tuning pipelines - Fal, Kling, Veo, Runway, and ElevenLabs as orchestrated external services - FP8/FP16 quantization, SageAttention, Lightning LoRAs, and other optimization techniques - FastAPI for communicating with our Azure backend ## What we're looking for - Strong Python skills: you write FastAPI services, async workers, inference scripts, or training scripts regularly - Practical diffusion model depth: samplers, schedulers, CFG, LoRA merging, VAEs, text encoders, and image/video generation workflows at the implementation level - ComfyUI fluency: you can build, debug, modify, and parameterize workflows programmatically - Serverless GPU experience: Modal, RunPod, Replicate, or similar - LoRA / fine-tuning experience: dataset preparation, training configs, evaluation, and iteration - Strong curiosity and pace: you actively follow the latest developments in generative media and can quickly turn new tools into working systems - Comfortable doing minor frontend/backend work when needed, especially with the help of AI coding tools - At least two years of full-time industry experience or equivalent practical experience ## Compensation Monthly compensation: USD 2,000–4,000 equivalent, depending on experience and seniority ## Nice to have - Interest in video, film, advertising, or creative production - Voice cloning or speech generation experience - Multi-model orchestration experience - Azure experience - Open-source contributions ## About the company View our work at: https://www.instagram.com/excury and https://excury.com/ Please do not apply before reading the full job description carefully.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗