Senior Machine Learning Engineer, Applied Science Data Frameworks
Core
Building foundational infrastructure for large-scale multimodal AI training and inference, including data loaders, feature enrichment pipelines, and dataset management systems.
Role type
Senior Machine Learning Engineer (Infrastructure)
Builds
Distributed training data loaders, feature enrichment pipelines, batch inference systems, and dataset registry systems for petabyte-scale foundation models.
Domain
AI/ML Infrastructure, Distributed Systems, Data Engineering
Deliverable
production ML models
Required skills
Distributed systems design, Python, System design, Data structures, Algorithms, Cloud platforms (AWS/Azure), Data platforms (Databricks/Spark), CI/CD, Containerization (Docker)
Preferred skills
ML frameworks (PyTorch/TensorFlow), MLOps practices, Batch inference architectures, Vector databases (OpenSearch/LanceDB)
Technologies
Apache Ray, Spark, DuckDB, Apache Arrow, Docker, OpenSearch, LanceDB
Responsibilities
Build and maintain distributed training data loaders for multi-source data ingestion; Implement feature enrichment pipelines and dataset registry systems; Develop batch inference pipelines for large-scale feature extraction; Optimize data pipeline performance for startup latency, throughput, and GPU utilization; Contribute to CI/CD infrastructure for ML systems; Write reusable framework components and SDKs.
Seniority
Senior, hands-on IC