CareerPlanSign in

Cloud Platform Engineer

Bengaluru, India💼 Full-time🗓 2026-07-08 → 2026-10-02

Core

Guardian of reliability, performance, and scalability for a full-stack generative AI inferencing service, bridging software development and operations to ensure exceptional uptime and low-latency response times.

Role type

Senior Cloud Site Reliability Engineer (AI Inferencing)

Builds

Production AI inferencing endpoints and the underlying cloud infrastructure for enterprise and government organizations.

Domain

Generative AI, High-Performance Computing, Cloud Infrastructure

Deliverable

production ML models

Required skills

Site Reliability Engineering, Cloud Infrastructure Management, Container Orchestration, Infrastructure as Code, CI/CD Pipeline Design, Monitoring and Observability, Capacity Planning, Incident Management, Python/Go/Java Programming

Preferred skills

Hybrid cloud/on-prem experience, ML/AI inferencing support, GPU-accelerated computing, Model serving frameworks (vLLM, SGLang, Ray), MLOps, Database and caching management

Technologies

AWS, GCP, Azure, Docker, Kubernetes, Terraform, Ansible, Prometheus, Grafana, Datadog, ELK Stack, Jenkins, GitHub Actions, ArgoCD, Redis, Memcached

Responsibilities

Manage production inferencing service availability, latency, and efficiency across multiple regions; Lead incident response and drive blameless post-mortems; Develop and maintain advanced monitoring and alerting systems; Design and implement auto-scaling policies; Manage cloud infrastructure using IaC; Build and improve CI/CD pipelines for model version deployment; Forecast infrastructure needs and optimize cloud costs.