AI Infrastructure Engineer
Core
Building and optimizing large-scale training infrastructure for Large Language Models (LLMs) and developing platform software to support AI/ML workflows.
Role type
Senior IC AI Infrastructure Engineer
Builds
Scalable AI infrastructure solutions, distributed training platforms, and containerized AI environments.
Domain
Artificial Intelligence / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
GPU programming, CUDA optimization, distributed systems, cloud computing, container technologies, Python, C++, PyTorch, Transformers
Preferred skills
Experience building large-scale distributed systems, optimizing neural network performance
Technologies
Docker, Kubernetes, CUDA, PyTorch, Transformers
Responsibilities
Designing and developing scalable AI infrastructure solutions for training and deploying large language models; Building and optimizing distributed training platforms using cutting-edge technologies; Implementing and maintaining containerized AI environments using Docker and Kubernetes; Optimizing CUDA kernels for maximum GPU utilization and performance; Developing platform software to support AI/ML workflows; Collaborating with AI researchers to implement efficient training and inference pipelines
Seniority
Senior, hands-on IC