DevOps Engineer
Core
Maintaining and developing a complex on-premise automation system deployed globally, ensuring operational continuity and resolving critical incidents at customer sites.
Role type
Senior DevOps Engineer (Infrastructure & Automation)
Builds
Automated troubleshooting mechanisms, health checks, observability solutions, and self-healing procedures for a multi-component system.
Domain
Enterprise Infrastructure, On-Premise Systems, GPU Computing
Deliverable
production ML models | infrastructure
Required skills
Linux (Ubuntu) system administration, Kubernetes cluster management, Docker containerization, Network troubleshooting, Storage (NFS) management, GPU/CUDA environment operations, RabbitMQ, PostgreSQL
Preferred skills
Designing automated self-healing mechanisms, Building observability dashboards and alerts, Creating runbooks and recovery procedures
Technologies
Kubernetes, Docker, Linux (Ubuntu), NFS, RabbitMQ, PostgreSQL, CUDA
Responsibilities
Diagnosing and resolving issues across Kubernetes, containers, OS, networking, and storage layers; Performing system deployments and upgrades at customer sites; Designing automated troubleshooting and validation mechanisms; Building health checks and observability solutions; Documenting incidents and root causes.
Seniority
Senior, hands-on IC