Director, Site Reliability Engineering
Core
Lead the Site Reliability Engineering organization to own reliability strategy, reduce failure frequency/impact, and ensure production readiness for Anduril's business and manufacturing operations.
Role type
Director, Site Reliability Engineering
Builds
CorpTech Platform (internal engineering force multiplier for corporate systems, Finance, Growth, and hardware enterprise)
Domain
Defense technology / Enterprise infrastructure / Manufacturing operations
Deliverable
production ML models | infrastructure
Required skills
Organizational leadership, distributed systems architecture, cloud infrastructure, container orchestration, networking, storage, observability, incident management, SLO definition, capacity planning, disaster recovery, risk assessment, strategic planning
Preferred skills
Scaling SRE in hyper-growth environments, federated engineering models, AI-enabled production systems, enterprise systems (ERP, MES, WMS, CRM)
Technologies
Kubernetes, AWS/Azure/GCP, Prometheus, Grafana, Terraform, Jenkins, Python, Go
Responsibilities
Build and lead the SRE organization (hiring, coaching, succession planning), set portfolio-wide reliability strategy and roadmap, establish shared accountability and service ownership with software engineering, define observability and incident management practices, lead response to critical incidents, reduce operational toil through automation and architectural changes, create healthy on-call systems, establish reliability practices for AI-enabled systems.
Seniority
Director, strategic leadership & hands-on IC