Principal Speech Data Linguist
Core
Define linguistic standards, quality frameworks, and human-in-the-loop workflows for high-volume speech segmentation and transcription to support AI model training and evaluation.
Role type
Principal Speech Data Linguist (applied expert)
Builds
Linguistic standards, quality rubrics, and annotation workflows for speech and audio models
Domain
Artificial Intelligence / Speech & Audio Technology (via careerplan.io/jobs/4418503009-principal-speech-data-linguist-at-innodatainc)
Deliverable
production ML models
Required skills
Linguistic judgment and standard-setting, phonetics and phonology, acoustic analysis (IPA, spectrograms, formants), audio segmentation and diarization, multilingual capability (accented/dialectal speech), Python scripting for data pipelines, quality lifecycle management
Preferred skills
Advanced degree in linguistics/phonetics, experience with low-resource languages, responsible AI considerations
Technologies
Praat, Montreal Forced Aligner, ELAN, Whisper, AssemblyAI, Deepgram, Rev, Speechmatics
Responsibilities
Define transcription/segmentation standards and style guides; establish quality frameworks (rubrics, error taxonomies, inter-annotator agreement); manage end-to-end quality lifecycle; design human-in-the-loop workflows; adjudicate linguistically hard cases; partner with research scientists on model specifications; train and mentor expert transcribers
Seniority
Principal, hands-on IC with strategy & mentorship
