AI Operations Engineer
Core
Design, build, and operate end-to-end infrastructure for production AI and LLM systems, ensuring reliability, observability, and data quality.
Role type
Machine Learning Operations / Language Model Operations Engineer
Builds
Production-grade data platforms, ingestion pipelines, observability tools, and evaluation frameworks for ML/LLM solutions.
Domain
Artificial Intelligence / Large Language Models / Data Engineering
Deliverable
production ML models
Required skills
Python engineering, data pipeline design, ML/LLM evaluation system development, LLM observability and tracing, PII redaction implementation, distributed data processing, system reliability engineering, failure mode analysis
Preferred skills
Multi-tenant/SaaS architecture, cloud ecosystem expertise (Azure/AWS), infrastructure trade-off decision making, vector databases and embedding techniques, data versioning tools (LakeFS/DVC/Delta Lake), GDPR compliance workflows, CI/CD and DevOps practices, frontend dashboard development (TypeScript)
Technologies
Python, Azure, AWS, LakeFS, DVC, Delta Lake, TypeScript, GitHub
Responsibilities
Design scalable data ingestion pipelines for production ML/LLM; Build data processing workflows including storage, classification, and slicing; Develop observability and debugging tools for production AI; Implement evaluation and monitoring frameworks including autoraters and regression detection; Build automated triage systems for production failures; Implement PII redaction and data governance mechanisms; Design LLM evaluation mining workflows; Implement alerting systems for model/prompt regressions; Evaluate and select tooling and hosting strategies for ML/LLM platforms; Own operational reliability of the entire ML/LLM data and evaluation pipeline
Seniority
Mid-to-Senior, hands-on IC