Senior AI Infrastructure Engineer
Core
Build, scale, and optimize the end-to-end machine learning platform and MLOps tooling that powers autonomous defense systems.
Role type
Senior AI Infrastructure Engineer
Builds
Scalable training, orchestration, and experimentation infrastructure for cloud and air-gapped edge environments
Domain
Defense technology / Autonomous systems / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
Python, Go, C++, container orchestration (Docker, Kubernetes), distributed training frameworks (PyTorch Distributed, Ray, Slurm, Megatron-LM), distributed data pipelines, end-to-end project ownership
Preferred skills
GPU/accelerator workload profiling, multi-tenant cluster management, production ML observability frameworks, RLHF/DPO pipeline support
Technologies
Kubernetes, Docker, PyTorch, Ray, Slurm, Megatron-LM
Responsibilities
Build and maintain scalable training and orchestration infrastructure; Develop tooling for experiment tracking, automated profiling, and hyperparameter tuning; Implement and scale robust ETL pipelines for multi-modal data; Deploy high-throughput, low-latency model serving frameworks; Develop CI/CD pipelines for ML models with automated testing and safe rollout strategies; Implement pipelines for model evaluation and reinforcement learning alignment loops; Mentor peers and conduct design/code reviews
Seniority
Senior, hands-on IC