Senior AI Engineer
Core
Design, build, and maintain the infrastructure and toolchains for large-scale distributed training and inference of Machine Learning models, specifically focusing on GPU clusters and LLMs.
Role type
Senior AI Infrastructure Engineer (LLM Training & Inference)
Builds
High-performance GPU training clusters, model serving architecture, and automated ML development pipelines.
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Distributed training systems (Horovod, DeepSpeed, PyTorch Distributed, Ray), Tensor/model parallelism, Cloud-native infrastructure (Kubernetes, Docker, Slurm), Systems programming (Python, Bash, C++, Go), Performance profiling and optimization, Automation (Prometheus, Grafana, Weights & Biases)
Preferred skills
Edge device deployment, Serverless architectures, Auto-scaling strategies, Model caching mechanisms
Technologies
Horovod, DeepSpeed, PyTorch Distributed, Ray, Kubernetes, Docker, Slurm, Prometheus, Grafana, Weights & Biases, CUDA, Triton, NCCL, gRPC, DALI, tf.data
Responsibilities
Develop and maintain high-performance LLM training GPU infrastructure and clusters; Optimize GPU utilization and implement fault-tolerant distributed training strategies; Build automated pipelines for data preprocessing, feature engineering, and model deployment; Design systems for dynamic resource allocation and model loading/unloading for inference; Develop dashboards and alerting mechanisms for real-time model performance monitoring; Provide technical support and troubleshooting for training and inference workloads.
Seniority
Senior, hands-on IC