Lead SysOps Engineer
Core
Design, deploy, and maintain large-scale Linux/Windows server clusters, HPC environments, and cloud infrastructure for complex technical workloads.
Role type
Principal DevOps/SysOps Engineer (Infrastructure & HPC)
Builds
Production-grade server clusters, containerized applications, and cloud-based solutions
Domain
High-Performance Computing (HPC), Cloud Infrastructure, System Administration
Deliverable
production ML models | infrastructure
Required skills
Linux server administration, HPC cluster management, Docker, Kubernetes, Terraform, Ansible, Bash scripting, Python scripting, CI/CD, GitOps, cloud platform management, firewall and load balancer configuration, network troubleshooting, cybersecurity best practices
Preferred skills
CKA/CKAD/CKS certification, additional automation tool proficiency, open-source community participation
Technologies
Docker, Kubernetes, Terraform, Ansible, SLURM, Run:AI, Infiniband, Ethernet, AWS, Azure, Google Cloud
Responsibilities
Deploy and maintain Linux/Windows servers across physical and virtual environments; manage HPC clusters and job scheduling systems; implement and optimize Docker/Kubernetes in large-scale environments; automate infrastructure using IaC tools; troubleshoot complex system and networking issues; configure firewalls, load balancers, and VPNs; apply cybersecurity best practices; collaborate with development teams on the software development lifecycle
Seniority
Principal, hands-on IC with strategic oversight
