HPC & MLOps Engineer
Core
Design, deploy, and maintain high-performance computing (HPC) and MLOps systems for clients, ensuring reliable and scalable AI infrastructure.
Role type
Senior IC HPC & MLOps Engineer
Builds
Cloud and on-premise HPC clusters, automated deployment pipelines, and MLOps workflows for AI/ML teams
Domain
AI Infrastructure / High-Performance Computing / Cloud Systems
Deliverable
production ML models | infrastructure
Required skills
HPC operations, Slurm/PBS/OpenPBS, Python (asyncio), Bash, Linux system administration, cloud platforms (AWS/GCP/Azure/OCI), Lustre/Ceph/NFS/object storage, GPU scheduling, RDMA/InfiniBand networking
Preferred skills
CUDA/NCCL profiling, MPI optimization, kernel/sysctl tuning, C/C++ systems programming, cost optimization in cloud bursting scenarios
Technologies
OpenMPI, CUDA, TensorFlow, PyTorch, Lustre, BeeGFS, Ceph, NFS, object storage, Slurm, PBS Pro, OpenPBS, AWS, GCP, Azure, OCI, Bash, Python, asyncio, aiohttp, boto3, google-cloud, azure-sdk, OCI SDK
Responsibilities
Configure, deploy, monitor, and maintain C-Gen.AI Cluster solutions; manage cloud and on-premise deployments with optimal job scheduling; troubleshoot and optimize HPC library stacks and parallel file systems; develop automated deployment scripts and workflows; instrument systems to build dashboards and alerts for incident response; collaborate on advanced HPC development projects involving performance tuning and GPU parallelization
Seniority
Senior, hands-on IC