CareerPlanSign in

Senior Site Reliability Engineer

California - San Francisco💼 Full-time🗓 2026-07-15 → 2026-09-26

Core

Senior SRE leading operational resilience, incident response, and AI-powered automation for Salesforce's global cloud services.

Role type

Senior IC Site Reliability Engineer (AI/ML Ops)

Builds

Production-grade observability solutions, self-healing systems, and AI-driven operational tooling for Salesforce's cloud infrastructure.

Domain

Cloud Infrastructure / SRE / AI Operations

Deliverable

production ML models | infrastructure

Required skills

Python, Go, Kubernetes, Linux/Unix internals, distributed systems, incident management, SRE principles (SLIs/SLOs), workflow orchestration (Temporal/Airflow/Argo), AI/ML for ops (anomaly detection, predictive analysis, prompt engineering)

Preferred skills

Experience with MCP-based agents, building durable automation pipelines, mentoring junior engineers

Technologies

Docker, Kubernetes, Temporal, Airflow, Argo Workflows, Grafana, Prometheus, ELK, Splunk, Datadog

Responsibilities

Lead incident detection, response, and resolution; architect and build production-grade observability solutions; design and implement AI/ML-powered operations tools; drive optimization of system performance and cost-effectiveness; provide technical coaching to junior team members; collaborate with engineering teams to define and uphold SLAs/SLOs.

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.