Machine Learning Engineer, Speech LLM Training - San Francisco
Core
Building and training large-scale audio or speech models (SpeechLLMs, ASR, TTS) to create a hardware-software AI companion that amplifies human productivity.
Role type
Senior IC machine-learning engineer (speech LLM training)
Builds
Unified SpeechLLMs, advanced ASR, expressive TTS, and generative audio architectures for Plaud's hardware-software AI interface.
Domain
AI/ML, Speech Processing, Audio Engineering
Deliverable
production ML models
Required skills
Large-scale audio/speech model training, sequence modeling architecture design, distributed training cluster debugging, signal processing, acoustic representation modeling, PyTorch, JAX, GPU memory optimization, performance bottleneck resolution
Preferred skills
Text-based LLM pretraining/instruction tuning/RLHF, neural audio codec design, diffusion/flow matching/autoregressive architectures for speech, RL alignment techniques (RLHF/GRPO), end-to-end inference optimization (vLLM/TensorRT-LLM/SGLang), massive GPU cluster management (FSDP/DeepSpeed), Kubernetes orchestration
Technologies
PyTorch, JAX, vLLM, TensorRT-LLM, SGLang, FSDP, DeepSpeed, Kubernetes
Responsibilities
Design novel sequence modeling architectures, debug distributed training clusters, traverse the stack from signal processing to edge-device optimization, take ownership of ambiguous problems and drive them to production
Seniority
Senior, hands-on IC