Senior AI Engineer - Observability
Core
Design, build, and improve AI-powered product features (RAG, semantic search, agentic workflows) while defining evaluation pipelines and observability systems to ensure reliability, cost-effectiveness, and quality in production.
Role type
Senior AI Engineer (Observability & Evaluation)
Builds
AI-powered governance platform features, evaluation pipelines, golden datasets, and observability dashboards for production AI systems.
Domain
Enterprise Governance / AI Engineering / LLM Operations
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, C#/.NET, LLM application architecture, RAG systems, prompt engineering, retrieval quality, evaluation design, observability fundamentals, CI/CD, cloud environments
Preferred skills
LLMOps, Arize/Langfuse/LangSmith, Azure AI Search/Pinecone/Qdrant, Semantic Kernel/LangChain/LlamaIndex, LLM-as-judge, DeepEval, A/B testing, regulated environment experience
Technologies
Python, C#/.NET, Azure DevOps, Arize, Langfuse, LangSmith, W&B, OpenTelemetry, Azure AI Search, Pinecone, Qdrant, Weaviate, pgvector, Semantic Kernel, LangChain, LlamaIndex, AutoGen, DeepEval, Humanloop, Helicone
Responsibilities
Design and implement evaluation pipelines and scoring methodologies for LLM features; Build and maintain versioned golden datasets covering real-world use cases; Instrument LLM interactions and agent workflows using observability platforms; Analyze production traces to identify quality issues and cost spikes; Build reusable libraries and SDKs for AI instrumentation and evaluation; Coach product teams on AI quality standards and reliable release practices.
Seniority
Senior, hands-on IC