Lead Site Reliability Engineer, Enterprise Tech
Core
Lead SRE responsible for building reliability into the platform via production-grade software, automation, and AI-assisted workflows to ensure resiliency, security, and cost efficiency for J.P. Morgan's enterprise infrastructure.
Role type
Lead Site Reliability Engineer (Enterprise Tech)
Builds
Production-grade software, automation, control loops, self-healing tooling, and AI-integrated reliability workflows.
Domain
Financial Services / Enterprise Infrastructure / AI-assisted Operations
Deliverable
production ML models | product features | infrastructure
Required skills
Site Reliability Engineering (SRE), Python, Go, Java, C++, or Rust, Kubernetes, Terraform, CI/CD, Observability (Grafana, Prometheus, Splunk, Datadog, Dynatrace), SLI/SLO management, Incident Response, Nix fundamentals, Systems thinking, Security-first mindset, AI tool validation and guardrail establishment.
Preferred skills
Networking depth (routing, switching, packet analysis), Multi-domain infrastructure experience, AI prompt engineering and agent orchestration, Regulated environment experience, Engineering culture building.
Technologies
AI, CI/CD, Datadog, Dynatrace, Flow, Grafana, Kubernetes, Prometheus, Python, Rust, Splunk, Terraform
Responsibilities
Define and operationalize SLIs/SLOs with stakeholders, lead on-call and major incidents, drive toil reduction through automation, validate AI-assisted operational recommendations, establish guardrails for team AI usage, and own services end-to-end including reliability, performance, and cost.
Seniority
Lead, hands-on IC with mentorship responsibilities