CareerPlanSign in

Site Reliability Engineer II

India, Telangana, Hyderabad💼 Full-time🗓 2026-08-05 → 2026-09-26

Core

Ensure the reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads.

Role type

Senior Site Reliability Engineer (AI/High-Performance Computing)

Builds

AI infrastructure, containerized environments, automation tools for deployment and monitoring

Domain

Artificial Intelligence, High-Performance Computing, Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

incident management, root cause analysis, performance optimization, infrastructure automation, cloud platform management, distributed systems design, GPU/InfiniBand management, container orchestration, predictive analysis

Preferred skills

technical leadership, cross-functional collaboration, innovation in best practices

Technologies

Kubernetes, Docker, Azure, GPUs, InfiniBand

Responsibilities

Lead incident response and root cause analysis to minimize downtime. Identify and resolve bottlenecks in compute, storage, networking, and specialized hardware. Develop and maintain automation tools for deployment, monitoring, and management of AI infrastructure. Provide technical guidance in cloud and AI infrastructure technologies.

Seniority

Senior, hands-on IC with technical leadership

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.