Senior AI Infrastructure Engineer, LLM/AI Platforms
Core
Design, build, and deploy large-scale AI infrastructure for LLM training, fine-tuning, and inference to power next-generation AI-driven security products.
Role type
Senior AI Infrastructure Engineer (LLM Platforms)
Builds
Production-grade LLM training pipelines, inference serving systems, and data platforms for AI/ML workloads.
Domain
Cybersecurity / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
LLM infrastructure engineering, GPU cluster provisioning, MLOps, distributed training frameworks, inference optimization, data pipeline architecture, container orchestration, Infrastructure as Code
Preferred skills
Agentic workflow frameworks, distributed data processing, cybersecurity domain knowledge
Technologies
PyTorch, Ray, Megatron, JAX, vLLM, Triton, MLflow, Sagemaker, Vertex AI, CUDA, Docker, Kubernetes, Slurm, Airflow, Terraform, Ansible, AWS, GCP, OCI
Responsibilities
Provision and configure large GPU clusters for LLM workloads; Develop and optimize LLM model-serving infrastructure; Lead model lifecycle management including versioning and checkpointing; Design robust evaluation frameworks for model performance; Identify and address GPU utilization bottlenecks; Architect data platforms for RAG and AI Agentic Systems; Define and enforce MLOps/DataOps best practices; Collaborate with Data Scientists and Product Managers to transform prototypes into production services
Seniority
Senior, hands-on IC with mentorship responsibilities