CareerPlanGet AI match score →

Senior SRE Engineer

Taiwan💼 Full-time🗓 2026-06-05 → 2026-07-31

Core

Manage large-scale Linux environments, HPC clusters, multi-cloud infrastructure, and internal GenAI platforms.

Role type

Senior Site Reliability Engineer (Infrastructure & HPC)

Builds

Production-grade Linux systems, HPC compute/storage environments, multi-cloud infrastructure, CI/CD pipelines, and internal AI services.

Domain

High-Performance Computing, Cloud Infrastructure, Linux Systems Administration

Deliverable

production ML models | infrastructure

Required skills

Linux systems administration, Bash/Shell scripting, Python programming, Infrastructure as Code (Terraform/CDK), Container orchestration (Kubernetes/Docker), CI/CD pipeline design, Storage management (NAS/Lustre), HPC cluster operations, Root-cause analysis, Autonomous problem solving.

Preferred skills

HPC scheduling (Slurm), Parallel filesystems (Lustre/GPFS), Linux performance tuning (perf/eBPF), Database operations (MySQL/ClickHouse), Low-latency network tuning, LLM application development, Self-managed Kubernetes, GPU server operations.

Technologies

Linux, Bash, Ansible, Python, Slurm, Lustre, NAS, NFS, SMB, AWS, Alibaba Cloud, GCP, Terraform, AWS CDK, Docker, Kubernetes, EKS, GitLab, Jenkins, Airflow, LangChain, LangGraph, Bedrock, Elasticsearch, NVIDIA CUDA, nvidia-smi, DCGM.

Responsibilities

Troubleshoot and perform root-cause analysis on large-scale Linux environments; Write maintainable automation scripts using Bash, Ansible, and Python; Operate and monitor HPC clusters and storage systems; Manage multi-cloud environments and build containerized deployment workflows; Operate self-hosted CI/CD systems and design deployment pipelines; Build and maintain internal AI platforms and services.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗