Senior Site Reliability Engineer
Core
Senior SRE leading operational resilience, incident response, and AI-powered automation for Salesforce's global cloud services.
Role type
Senior IC Site Reliability Engineer (AI/ML Ops)
Builds
Production-grade observability solutions, self-healing systems, and AI-driven operational tooling for Salesforce's cloud infrastructure.
Domain
Cloud Infrastructure / SRE / AI Operations
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Kubernetes, Linux/Unix internals, distributed systems, incident management, SRE principles (SLIs/SLOs), workflow orchestration (Temporal/Airflow/Argo), AI/ML for ops (anomaly detection, predictive analysis, prompt engineering)
Preferred skills
Experience with MCP-based agents, building durable automation pipelines, mentoring junior engineers
Technologies
Docker, Kubernetes, Temporal, Airflow, Argo Workflows, Grafana, Prometheus, ELK, Splunk, Datadog
Responsibilities
Lead incident detection, response, and resolution; architect and build production-grade observability solutions; design and implement AI/ML-powered operations tools; drive optimization of system performance and cost-effectiveness; provide technical coaching to junior team members; collaborate with engineering teams to define and uphold SLAs/SLOs.
Seniority
Senior, hands-on IC with leadership responsibilities