AI-DNA Senior Site Reliability Engineer
Core
Senior SRE operating an AI-native community and social engagement platform for Fortune 100 brands, acting as first responder while building AI agents to automate incident resolution and operational tasks.
Role type
Senior Site Reliability Engineer (AI-native operations)
Builds
AI agents for pre-triage, change-validation, auto-healing, and RCA drafting; production-grade runbooks and guardrails.
Domain
SaaS / Social Engagement Platform / AI Operations
Deliverable
production ML models | infrastructure
Required skills
On-call incident management, root cause analysis, production change management, AWS multi-AZ/multi-account operations, AI agent development and tuning, automated remediation, runbook engineering, cost optimization.
Preferred skills
Advanced agentic tools (Claude Code, Codex, Warp), deep production scars, self-directed ownership, generalization of one-off fixes into scalable agents.
Technologies
AWS, Claude Code, Codex, Warp, custom agents
Responsibilities
Own shift and resolve customer-impacting incidents; build and maintain AI agents for operations; execute production changes with tested rollbacks; perform deep root cause analysis and prevent recurrence; generalize manual fixes into automated agents; document procedures for agent retrieval.
Seniority
Senior, hands-on IC