Senior Site Reliability Engineer
Core
Modernize application deployments by tackling technical debt on legacy Linux systems, building internal automation tools, and establishing observability stacks to ensure stability during cloud-native transitions.
Role type
Senior Site Reliability Engineer (Infrastructure Modernization & Automation)
Builds
Internal system management tools, observability stack, CI/CD pipelines, and infrastructure as code
Domain
Cloud Infrastructure, Linux Systems, DevOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux system administration, Python, Shell, Java, NodeJS, AWS, GCP, Azure, Docker, Kubernetes, Prometheus, Grafana, ELK stack, Terraform, Ansible, Git, CI/CD pipeline management
Preferred skills
Microservices architecture, distributed systems, security compliance frameworks, chaos engineering, open-source contributions
Technologies
Terraform, Ansible, Helm, GitLab CI/CD, GitHub Actions, Prometheus, Grafana, ELK stack, Docker, Kubernetes
Responsibilities
Design and implement OS/platform-level and application-level monitoring dashboards; Establish and maintain SLIs, SLOs, and error budgets; Build alerting systems for proactive issue resolution; Participate in on-call rotations and lead incident response efforts; Develop and optimize CI/CD pipelines for speed and resilience; Manage database reliability, performance, and scaling; Implement service discovery, load balancing, and network policies; Monitor and optimize cloud resource usage and costs.
Seniority
Senior, hands-on IC