Site Reliability Engineer, Litmus GCC
Core
Own the day-to-day reliability, security, and performance of an Azure-hosted industrial data platform (Litmus Unified Namespace and Edge Manager) for a strategic enterprise customer.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure)
Builds
Azure-hosted data pipelines, container orchestration environments, and identity federation systems for industrial AI applications.
Domain
Industrial AI, Edge Computing, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Azure cloud services, Kubernetes administration, Infrastructure as Code, scripting automation, networking fundamentals, incident management, SLA-driven operations
Preferred skills
Identity federation (OIDC/SAML), MQTT protocols, IoT/industrial environment support
Technologies
Azure AKS, Azure VNet/NSG, Azure Database, Terraform, Bicep, Python, Bash, PowerShell, Key Vault, Azure Monitor
Responsibilities
Provision and maintain cloud infrastructure; own end-to-end monitoring and alerting; drive on-call support and root cause analysis; implement security baselines and vulnerability management; manage networking and connectivity; build and maintain CI/CD pipelines; manage backup and disaster recovery; support customer validation and go-live activities.
Seniority
Senior, hands-on IC