Staff AI Infrastructure Engineer
Core
Architect, build, and scale the end-to-end machine learning platform powering autonomous defense systems (Fury, Barracuda) and Lattice OS.
Role type
Staff AI Infrastructure Engineer
Builds
End-to-end ML platform, MLOps tooling, and systems architecture for training, evaluating, hosting, and serving complex AI models (LLMs, CV, RL) in cloud and air-gapped edge networks.
Domain
Defense technology / Autonomous systems / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Go, C++, ML systems design, distributed computing, containerized deployments, GPU scheduling, distributed training frameworks, ETL pipelines, technical direction, mentoring
Preferred skills
Secure/air-gapped ML infrastructure, LLM/GenAI/RL training platforms, hardware performance profiling, multi-tenant ML platforms, next-gen AI accelerators (Trainium/TPUs/ASICs), ML observability
Technologies
Docker, Kubernetes, PyTorch Distributed, Ray, Slurm, Megatron-LM, AWS Trainium, Google TPUs
Responsibilities
Design and maintain foundational training, orchestration, and experimentation infrastructure; Build automated tools for hyperparameter tuning and model profiling; Design and scale high-performance ETL pipelines for terabytes of multi-modal data; Architect high-throughput, low-latency model serving frameworks for cloud and edge; Build CI/CD pipelines for ML models with automated validation and canary deployments; Work with researchers to design unified infrastructure standards.
Seniority
Staff, founding team, technical roadmap ownership & mentorship