Engineer II - Site Reliability (Hybrid, IND)
Core
Operate and evolve internal Temporal workflow orchestration infrastructure, a stateful distributed system running on Kubernetes, ensuring high availability and performance for engineering teams.
Role type
Junior Site Reliability Engineer (Infrastructure Operations)
Builds
Internal production workflow orchestration platform (Temporal) serving engineering teams
Domain
Cybersecurity / Cloud Infrastructure / Distributed Systems
Deliverable
infrastructure
Required skills
Kubernetes operations, Helm package management, Infrastructure-as-Code (Terraform/Ansible/FluxCD), Cloud platforms (AWS/GCP), Scripting (Bash/Python/Go), PostgreSQL basics, Incident response
Preferred skills
Workflow orchestration platforms (Airflow/Prefect), Distributed tracing, Go programming, Multi-region deployment patterns
Technologies
Kubernetes, Helm, FluxCD, Temporal, PostgreSQL, AWS, GCP, Terraform, Ansible, Bash, Python, Go
Responsibilities
Deploy updates and monitor cluster health, Automate operational tasks and scaling, Perform capacity planning and performance tuning, Build observability dashboards and alerting, Participate in on-call rotation and incident response, Troubleshoot deployment failures and connectivity issues
Seniority
Junior, growth-oriented IC