CareerPlanGet AI match score →

Senior Platform Reliability Engineer

Mexico City💼 Full-time🗓 2026-06-10 → 2026-07-31

Core

Designing and operating observability layers, automated remediation loops, and reliability controls for Gen AI platforms and internal agentic systems.

Role type

Senior Platform Reliability Engineer (SRE)

Builds

Telemetry systems, automated findings-to-fix workflows, and production readiness controls for AI platforms

Domain

Generative AI / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Python scripting, AWS cloud operations, observability stack design, infrastructure-as-code (Terraform/Pulumi), incident response, SLO/SLA management, CI/CD integration, API automation

Preferred skills

AI agent runtime experience, privacy-aware telemetry, security tooling (Wiz, CrowdStrike, GuardDuty), MCP integration knowledge

Technologies

Python, AWS, Terraform, Pulumi, CSPM tools, MCP

Responsibilities

Design observability layers with audit trails and traces; Build automated findings-to-fix loops; Implement reliability controls (alerting, kill-switches, rate limiting); Codify detections and policies as code; Review platform changes for hardening and production readiness; Partner with SOC teams on incident workflows

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗