CareerPlanGet AI match score →

Senior Site Reliability Engineer - Security

💼 Full-time🗓 2026-07-18 → 2026-07-21

Core

Designing and operating observability layers, automated remediation loops, and reliability controls for Gen AI platforms and internal agentic systems.

Role type

Senior Site Reliability Engineer (Security & Observability)

Builds

Telemetry pipelines, automated findings-to-fix workflows, and production hardening controls for AI systems.

Domain

Generative AI / Cloud Infrastructure / Security Automation

Deliverable

production ML models | infrastructure

Required skills

Python scripting, Infrastructure-as-Code (Terraform/Pulumi), observability (metrics/logs/traces), AWS logging and IAM diagnostics, CI/CD integration, incident response, SLO/SLA management

Preferred skills

Experience with AI agent runtimes, MCP/agent telemetry, privacy-aware telemetry, code review and testing habits

Technologies

Python, Terraform, Pulumi, AWS, Wiz, CrowdStrike, Orca, GuardDuty, WAF, RASP

Responsibilities

Design observability layers including audit trails and runtime visibility; Build automated remediation loops integrating CSPM signals; Implement reliability controls like kill-switch validation and drift detection; Codify detections and policies as code; Review platform changes for hardening and production readiness; Partner with SOC teams on incident workflows.

Seniority

Senior, hands-on IC

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗