Principal Administrator, Systems Management
Core
Principal SRE driving reliability, performance, and scalability of critical systems while championing AI-powered tooling adoption.
Role type
Principal individual contributor Site Reliability Engineer (SRE)
Builds
Production-grade AI/ML infrastructure pipelines, observability platforms, automation frameworks, and resilient cloud-native services
Domain
Pharmaceutical industry, Cloud Infrastructure, AI/ML Operations, Site Reliability Engineering
Deliverable
production ML models | infrastructure | product features
Required skills
Site Reliability Engineering (SRE), Cloud Platform Architecture, AI/ML Platform Integration, Network Engineering, Infrastructure as Code, Observability Strategy, Automation Development, Incident Management, Technical Mentorship, Cross-functional Leadership
Preferred skills
LLM-based workflow deployment, FinOps principles, Regulated industry compliance knowledge, Kubernetes orchestration, Cloud Architecture Certifications
Technologies
Azure AI Foundry, Azure OpenAI Service, Azure Machine Learning, Python, Go, Terraform, Bicep, Azure DevOps, Datadog, Grafana, Prometheus, TCP/IP, SD-WAN, Kubernetes
Responsibilities
Define and enforce SLOs/SLIs/error budgets; Lead major incident response and post-incident reviews; Architect and maintain AI/ML infrastructure pipelines; Design advanced observability strategies; Write production-quality automation code; Mentor engineers on SRE principles; Collaborate with data science teams on AI integration
Seniority
Principal, hands-on IC with strategic influence