Senior SRE Engineer
Core
Manage large-scale Linux environments, HPC clusters, multi-cloud infrastructure, and internal GenAI platforms.
Role type
Senior Site Reliability Engineer (Infrastructure & HPC)
Builds
Production-grade Linux systems, HPC compute/storage environments, multi-cloud infrastructure, CI/CD pipelines, and internal AI services.
Domain
High-Performance Computing, Cloud Infrastructure, Linux Systems Administration
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, Bash/Shell scripting, Python programming, Infrastructure as Code (Terraform/CDK), Container orchestration (Kubernetes/Docker), CI/CD pipeline design, Storage management (NAS/Lustre), HPC cluster operations, Root-cause analysis, Autonomous problem solving.
Preferred skills
HPC scheduling (Slurm), Parallel filesystems (Lustre/GPFS), Linux performance tuning (perf/eBPF), Database operations (MySQL/ClickHouse), Low-latency network tuning, LLM application development, Self-managed Kubernetes, GPU server operations.
Technologies
Linux, Bash, Ansible, Python, Slurm, Lustre, NAS, NFS, SMB, AWS, Alibaba Cloud, GCP, Terraform, AWS CDK, Docker, Kubernetes, EKS, GitLab, Jenkins, Airflow, LangChain, LangGraph, Bedrock, Elasticsearch, NVIDIA CUDA, nvidia-smi, DCGM.
Responsibilities
Troubleshoot and perform root-cause analysis on large-scale Linux environments; Write maintainable automation scripts using Bash, Ansible, and Python; Operate and monitor HPC clusters and storage systems; Manage multi-cloud environments and build containerized deployment workflows; Operate self-hosted CI/CD systems and design deployment pipelines; Build and maintain internal AI platforms and services.
Seniority
Senior, hands-on IC