Site Reliability Engineer- AI Enablement
Core
Enable engineering teams to adopt AI responsibly by evaluating architectures, guiding tooling, and ensuring operational readiness for AI systems.
Role type
Senior IC Site Reliability Engineer (AI Enablement)
Builds
Reliable, observable, and governance-aligned AI-powered services for healthcare organizations
Domain
Healthcare IT + AI/LLM Systems
Deliverable
production ML models
Required skills
AI system solutioning, LLM API integration, agentic/RAG frameworks, SRE principles, observability, cloud infrastructure (Azure/AWS), container orchestration, AI governance
Preferred skills
Software engineering background, healthcare IT compliance, AI evaluation/red-teaming, rules engines, Databricks, observability tooling
Technologies
Azure AI Foundry, Anthropic Claude, LangChain, LlamaIndex, Semantic Kernel, Docker, Kubernetes, Datadog, Grafana, OpenTelemetry, Databricks
Responsibilities
Train engineering teams on AI workflows and tooling; evaluate AI system designs for reliability and governance; advise on SLOs and failure modes; develop internal standards and reference architectures; participate in incident calls as an AI subject matter expert; maintain documentation of AI patterns.
Seniority
Senior, hands-on IC