Site Reliability Engineer I
Core
Support and improve foundational networking, compute, Kubernetes, and ingress/traffic-management infrastructure to ensure platform reliability and scalability.
Role type
Junior Site Reliability Engineer (IC)
Builds
Production networking, compute, and Kubernetes infrastructure for PagerDuty's platform
Domain
Cloud-native infrastructure (AWS/GCP/Azure) and container orchestration
Deliverable
production ML models | product features | infrastructure
Required skills
Linux system administration, networking fundamentals (load balancing, DNS, TLS, ingress), container orchestration (EKS/Kubernetes), cloud-native infrastructure (AWS/GCP/Azure), programming (Python/Ruby/Go), Infrastructure as Code (Terraform/CloudFormation)
Preferred skills
AWS cloud networking (VPCs, subnets, routing, security groups, load balancers), production Kubernetes operations (cluster upgrades, networking, ingress), observability platforms (Datadog, New Relic, SumoLogic, Splunk, Prometheus, Grafana), service meshes (Envoy, Istio, NGINX)
Responsibilities
Harden existing systems and support rollout of new infrastructure capabilities, monitor system health via metrics/logs/alerts, participate in 24/7 on-call rotations for incident response, participate in team planning and progress communication
Seniority
Junior, hands-on IC
