Senior Staff Engineer, ML Ops (R4941)
Core
Design and build a Kubernetes-native AI Factory Reference Architecture for developing, training, evaluating, and deploying next-generation AI systems for defense autonomy.
Role type
Senior Staff Engineer, MLOps Platform
Builds
Kubernetes-native platform for distributed AI training, simulation, evaluation, and deployment
Domain
Defense technology, AI infrastructure, distributed systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Python, Golang, distributed training, GPU infrastructure, Terraform, Helm, PyTorch, Hugging Face Transformers
Preferred skills
Ray, KAI, Slurm, reinforcement learning, edge deployment, observability (OpenTelemetry, Prometheus, Grafana)
Technologies
Kubernetes, PyTorch, Hugging Face, Terraform, Helm, Ray, KAI, Slurm, OpenTelemetry, Prometheus, Grafana
Responsibilities
Lead design and implementation of AI Factory Reference Architecture; Partner with ML researchers to support evolving training workflows; Design self-service AI development workflows; Build infrastructure for distributed training and inference; Optimize shared GPU infrastructure; Build capabilities for data and model lifecycle management; Develop deployment solutions using Infrastructure as Code; Evaluate emerging AI infrastructure technologies.
Seniority
Senior Staff, hands-on IC with architectural leadership