CareerPlanSign in

Senior Site Reliability Engineer

India, Telangana, Hyderabad💼 Full-time🗓 2026-08-05 → 2026-09-26

Core

Ensure reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads, leading incident response and performance optimization.

Role type

Senior Site Reliability Engineer (AI/Cloud Infrastructure)

Builds

AI infrastructure services, containerized environments, and automation tools for deployment and monitoring.

Domain

Artificial Intelligence, High-Performance Computing, Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Incident management, root cause analysis, performance optimization, infrastructure automation, Kubernetes, Docker, cloud platforms, distributed systems, GPU management, InfiniBand, predictive analysis

Preferred skills

Publications, certifications in cloud or AI infrastructure

Responsibilities

Lead incident response and root cause analysis, identify and resolve bottlenecks in compute/storage/networking, develop automation tools for deployment and monitoring, provide technical guidance on cloud and AI technologies, advocate for service excellence, research emerging AI infrastructure technologies

Seniority

Senior, hands-on IC with technical leadership

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.