Staff Data Engineer (Audio/ML)
Core
Design, implement, and maintain scalable data pipelines for ingesting, preprocessing, and transforming large-scale audio datasets to support AI/ML model training and evaluation for immersive audio applications.
Role type
Staff Data Engineer (Audio/ML)
Builds
Robust data pipelines for AI/ML research in speech processing, style transfer, and source separation
Domain
Audio engineering and machine learning
Deliverable
production ML models
Required skills
Python, Pandas, NumPy, PyTorch, Librosa, FFmpeg, SoX, Apache Spark, Airflow, Docker, Kubernetes, AWS S3, Redshift, Google BigQuery
Preferred skills
Active learning, distributed training pipelines, open-source contributions, data visualization tools, machine learning principles
Technologies
GitLab, Tableau, Matplotlib
Responsibilities
Design and maintain automated data pipelines for audio dataset ingestion and transformation; Develop advanced preprocessing techniques for immersive and multichannel audio formats; Automate data cleaning, normalization, and augmentation processes; Integrate external datasets and APIs; Monitor and optimize pipeline performance; Create tools and workflows for annotating and curating datasets; Perform exploratory data analysis to validate dataset quality
Seniority
Staff, hands-on IC with strategic oversight