CareerPlanSign in

Principal Cloud Operations Engineer (10166)

San Jose, California, United States💼 Full-time💰 $180,000–$260,000🗓 2026-01-13 → 2026-09-25

Core

Lead cloud infrastructure implementation and provide technical leadership in cloud architecture, operational excellence, and cost optimization for large-scale production environments.

Role type

Principal Cloud Operations Engineer

Builds

Public/private/local cloud solutions, Kubernetes-based microservices deployment platforms, and automation tooling.

Domain

Cloud Infrastructure & DevOps

Deliverable

production ML models | product features | infrastructure

Required skills

Cloud architecture design, operational excellence, cost optimization, AI integration, cloud security, root cause analysis, automation development, distributed team collaboration, performance analysis

Preferred skills

Cloud security and compliance implementation

Technologies

AWS, Google Cloud, Azure, Kubernetes, Docker, Terraform, Helm, ArgoCD, SQL, NoSQL, Elasticsearch, PostgreSQL, Redis, Ignite, Flink, Kafka, RabbitMQ, Nagios, Grafana, Prometheus

Responsibilities

Provide technical leadership in cloud architecture, operational excellence, reliability, and cost optimization across large-scale production environments. Stay current with industry trends and leverage AI technologies and cloud service provider platforms to improve operational efficiency, scalability, security, and resiliency. Design and ensure secure, reliable, and high-performance communication across multiple regions and cloud service providers. Configure, tune, and operate middleware services, including SQL and NoSQL databases, messaging and streaming platforms, and related infrastructure components. Evaluate, recommend, and lead the adoption of CloudOps and DevOps tools, platforms, and automation solutions. Troubleshoot complex production infrastructure and application issues, providing deep technical expertise and hands-on support when required. Drive root cause analysis (RCA), implement corrective actions, and establish preventive measures to avoid recurrence. Collaborate closely with engineering cloud architects in system design discussions, architecture reviews, and whiteboard sessions. Partner with Development, QA, SRE, and external service providers or carriers to resolve issues and improve system reliability. Design, implement, and evolve deployment automation platforms for Kubernetes-based microservices. Improve service availability, performance, and scalability through automation, tooling, capacity planning, and process improvements. Analyze system and service performance, identify bottlenecks, and deliver actionable recommendations to improve efficiency and resilience.

Seniority

Principal, hands-on IC with strategic leadership

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.