CareerPlanGet AI match score →
Onsite or remote • Hyderabad+2🌐 Remote💼 Full-time🗓 2026-06-25

Core

Design, implement, and maintain scalable CI/CD pipelines and cloud infrastructure for distributed systems and AI-driven platforms.

Role type

Senior DevOps Engineer (AI/ML infrastructure)

Builds

Scalable CI/CD pipelines, containerized microservices, GPU orchestration environments, and AI/ML model-serving infrastructure

Domain

Cloud Infrastructure, DevOps, AI/ML Engineering

Deliverable

production ML models | infrastructure

Required skills

CI/CD pipeline architecture, AWS/GCP cloud infrastructure, Docker & Kubernetes, Infrastructure as Code (Terraform, Ansible), Python/Bash/Go scripting, observability (Prometheus, Grafana, ELK), AI/ML pipeline support (MLflow, Kubeflow, SageMaker)

Preferred skills

GPU workload management, serverless architectures, GitOps tools, AIOps platforms

Technologies

GitHub Actions, Terraform, Ansible, CloudFormation, Docker, Kubernetes, EKS, GKE, Prometheus, Grafana, CloudWatch, ELK, MLflow, Kubeflow, SageMaker, Vertex AI, Lambda, Cloud Run, Argo CD, Flux, Dynatrace, New Relic

Responsibilities

Design and maintain CI/CD pipelines for Prerel, QA, and Production; Automate provisioning and configuration using IaC tools; Architect and maintain cloud infrastructure (AWS/GCP); Lead modernization initiatives for containerization and microservices; Implement high availability setups and DR strategies; Support AI/ML model training and deployment pipelines; Manage GPU orchestration and scalable model-serving environments; Implement centralized monitoring, logging, and alerting systems; Ensure cloud security best practices and compliance with HIPAA, SOC2, ISO 27001

Seniority

Senior, hands-on IC

Rewrite
## About the role Ekshvaku Tech Innovations is looking for a highly experienced Senior DevOps Engineer who excels at building scalable, automated, and resilient infrastructure for modern distributed systems and AI-driven platforms. You will lead DevOps strategy, modernize our infrastructure, and ensure high performance across multi-environment deployments. This role requires strong hands-on expertise, architectural thinking, excellent documentation discipline, and the ability to collaborate with cross-functional engineering teams. ## Responsibilities - Design, implement, and maintain scalable CI/CD pipelines for Prerel, QA, and Production. - Continuously improve build, release, and deployment workflows. - Reduce manual ops through automation and scripting. - Automate provisioning and configuration using Terraform, Ansible, and similar IaC tools. - Architect, deploy, and maintain cloud infrastructure (AWS/GCP). - Lead modernization initiatives: containerization, orchestration, microservices optimization. - Implement high availability setups, DR strategies, and cost-optimized infrastructure. - Support AI/ML model training and deployment pipelines. - Manage GPU orchestration and scalable model-serving environments. - Work closely with AI and Data Engineering teams. - Implement centralized monitoring, logging, and alerting systems. - Improve system reliability and incident response times. - Ensure cloud security best practices (IAM, network security, zero-trust). - Contribute to compliance efforts for HIPAA, SOC2, ISO 27001. - Maintain clear documentation, runbooks, and architecture diagrams. - Collaborate with Engineering, Product, QA, and AI teams. - Evaluate new DevOps tools and AI-powered automation capabilities. ## Requirements - 8+ years in DevOps or Cloud Infrastructure roles. - Strong expertise in GitHub Actions and CI/CD pipeline architecture. - Deep understanding of AWS or GCP (VPC, IAM, security, autoscaling). - Production-level experience with Docker & Kubernetes (EKS/GKE). - Proficiency in scripting: Python, Bash, or Go. - Strong experience with Terraform, Ansible, CloudFormation. - Solid understanding of monitoring tools (Prometheus, Grafana, CloudWatch, ELK). - Experience supporting AI/ML pipelines (MLflow, Kubeflow, SageMaker, Vertex AI). - Knowledge of security best practices & compliance frameworks. - Excellent communication and documentation skills. ## Nice to have - Experience with GPU workloads and ML orchestration. - Knowledge of event-driven or serverless architectures (Lambda, Cloud Run). - Exposure to GitOps tools (Argo CD, Flux). - Background in AIOps (Dynatrace Davis, New Relic AI). - AWS DevOps Engineer Professional or equivalent certification. ## What we offer - Location: Remote (Must be available to work in CST Timezone) - Experience Required: 8+ Years - Department: Engineering - Employment Type: Full-Time ## About the company Success Indicators - 90%+ reduction in manual deployment tasks within 6 months. - Fully standardized infrastructure documentation across all environments. - Reduced downtime and faster incident resolution through observability. - Improved developer velocity and increased deployment frequency. - Adoption of AI-assisted automation and predictive monitoring.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗