Senior SysOps Engineer
Core
Design, deploy, and maintain large-scale Linux/Windows server clusters, HPC environments, and cloud infrastructure for software development and high-performance computing workloads.
Role type
Principal DevOps/SysOps Engineer
Builds
Production-grade server clusters, containerized applications, CI/CD pipelines, and cloud-based solutions
Domain
High-Performance Computing (HPC), Cloud Infrastructure, System Administration
Deliverable
production ML models | infrastructure
Required skills
Linux server administration, HPC cluster management (SLURM/Run:AI), Docker, Kubernetes, Terraform, Ansible, Bash scripting, Python scripting, cloud platform management (AWS/Azure/Google), firewall and load balancer configuration, cybersecurity best practices, monitoring tool implementation
Preferred skills
CKA/CKAD/CKS certification, additional automation tool proficiency, open-source community participation
Technologies
Docker, Kubernetes, Terraform, Ansible, SLURM, Run:AI, Infiniband, Ethernet, Bash, Python
Responsibilities
Deploy and maintain Linux/Windows servers in physical and virtual environments; manage HPC clusters and job scheduling systems; implement and optimize Docker and Kubernetes in large-scale environments; configure CI/CD pipelines and GitOps workflows; troubleshoot complex system and networking issues in large-scale clusters; manage firewalls, load balancers, and VPNs; apply cybersecurity best practices to enhance system security; utilize infrastructure-as-code tools for reproducible deployments; collaborate with development teams on the software development lifecycle
Seniority
Principal, hands-on IC with strategic oversight