Site Reliability Engineer II (SRE)
Core
Own assigned infrastructure and software reliability problem statements, develop and maintain AWS, Kubernetes, and observability tools, and participate in on-call schedules.
Role type
Site Reliability Engineer II (SRE)
Builds
AWS accounts, networking, Kubernetes clusters, observability tools, and CI/CD pipelines
Domain
Cloud Infrastructure & DevOps
Deliverable
production ML models | infrastructure
Required skills
Python, JavaScript, Unix Shell, AWS, Terraform, Kubernetes, CI/CD pipelines, Docker
Responsibilities
Write, improve, and document efficient code and existing systems; Develop, configure, and maintain AWS accounts, networking, Kubernetes clusters, and other underlying platforms; Develop, configure, and maintain observability and monitoring tools including Coralogix and Sentry; Develop, configure, and maintain software development tools including GitHub Actions runners and Argo CD; Participate in an operational on-call schedule.
Seniority
Mid-level, hands-on IC