Site Reliability Engineering 2
Core
Partner with service engineering and product teams to build reliable self-service experiences and improve the quality, reliability, and safety of AI-generated responses through evaluation, monitoring, and feedback loops.
Role type
Site Reliability Engineer (AI/LLM focus)
Builds
Reliable self-service experiences for engineers and customers; operational health of AI services
Domain
Artificial Intelligence / Large Language Models / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
C, C++, C#, Java, JavaScript, Python, Retrieval-Augmented Generation (RAG), prompt engineering, model output evaluation, Kusto Query Language (KQL), telemetry, monitoring, live-site operations
Preferred skills
Experience building AI or large-language-model-powered applications
Technologies
Kusto / Azure Data Explorer, KQL
Responsibilities
Identify high-value engineer and customer scenarios for self-service; Improve quality and safety of AI responses via evaluation and monitoring; Contribute to operational health via telemetry and alerting; Write well-tested, maintainable code; Collaborate on design discussions and code reviews
Seniority
Mid-Senior, hands-on IC