CareerPlanSign in

Operations Engineer

Bengaluru💼 Full-time🗓 2026-09-22 → 2026-09-25

Core

Build and maintain the reliability of infrastructure engineering tooling estate including observability platforms, IaC, CI/CD pipelines, ITSM platforms, and AI-augmented operations tooling in a regulated financial services environment.

Role type

Senior Site Reliability Engineer (Tools & Platforms)

Builds

Automated remediation workflows, SLI/SLO definitions, GitOps pipelines, and AI-augmented operations tooling for internal developer platforms.

Domain

Financial Services / Platform Engineering & AI Operations

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Splunk Enterprise Architecture, OpenTelemetry, Terraform, Python, Incident Management, Capacity Planning, LLMOps, RAG Platform Reliability, Vector Databases, CI/CD Pipelines, Secret Management

Preferred skills

Chaos Engineering, eBPF-based observability, Model Serving Infrastructure, Financial Services Domain Experience

Technologies

Splunk, OpenTelemetry, ELK, Terraform, Ansible, GitHub Actions, ArgoCD, HashiCorp Vault, ServiceNow, Backstage, LangChain, LlamaIndex, CrewAI, Pinecone, Weaviate, ChromaDB, AWS, GCP

Responsibilities

Define and maintain SLIs/SLOs for observability and AI tooling pipelines; Engineer auto-remediation for common platform failures; Implement and govern OpenTelemetry instrumentation standards; Drive observability-as-code adoption via GitOps; Perform capacity planning and performance analysis for observability platforms; Lead blameless post-mortems for platform and AI tooling failures; Monitor model API health and prompt execution success rates for LLMOps; Ensure RAG knowledge base availability and retrieval latency SLOs.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.