Senior Site Reliability Engineer
Core
Run production environments, build infrastructure software, and ensure reliability for large distributed software applications.
Role type
Senior Site Reliability Engineer (IC)
Builds
Platform infrastructure, automation systems, and distributed software services
Domain
Cloud computing, DevOps, and distributed systems
Deliverable
production ML models | product features | infrastructure
Required skills
Systems engineering, automation, cloud platform expertise, infrastructure as code, CI/CD, containerization, microservices, monitoring and observability, incident management
Preferred skills
Large Kubernetes cluster administration, Grafana Observability Suite, configuration management tools, cloud certifications
Technologies
AWS, EC2, ECS, Lambda, DynamoDB, CloudFormation, Terraform, Jenkins, GitLab CI/CD, CircleCI, Docker, Kubernetes, Prometheus, Grafana, ELK stack, Cloudwatch, Splunk, Datadog, Pagerduty, Rundeck, Ansible, Puppet, Chef
Responsibilities
Monitor system availability and health, build software to manage platform infrastructure, optimize system performance, provide operational support for distributed applications, analyze metrics for performance tuning, partner with development on testing and releases, participate in system design and capacity planning, create sustainable systems through automation, balance feature speed with reliability objectives, manage incidents and postmortems
Seniority
Mid-Senior, hands-on IC