Principal Machine Learning Infrastructure Engineer
Core
Design and operate distributed training infrastructure and model serving pipelines for Large Physics Models (neural operators) to enable AI-driven simulation in engineering and manufacturing.
Role type
Principal Machine Learning Infrastructure Engineer
Builds
Distributed training clusters, model serving infrastructure, experiment tracking systems, and deployment pipelines for physics simulation models.
Domain
Deep-tech / AI for Engineering / Numerical Physics / Simulation
Deliverable
production ML models
Required skills
Distributed training (NCCL, FSDP, DDP, pipeline parallelism), Linux systems administration, Kubernetes, SLURM, Python, PyTorch, Cloud GPU infrastructure, Data I/O optimization, CI/CD
Preferred skills
Geometric deep learning, Neural operators, HPC for simulation (CFD/FEA), Model packaging for customer environments, Experiment tracking (Weights & Biases, MLflow), Observability (Prometheus, Grafana)
Technologies
NVIDIA DGX B200, NVLink, InfiniBand, Kubernetes, SLURM, PyTorch, CoreWeave, Prometheus, Grafana, W&B, MLflow
Responsibilities
Design and operate distributed training infrastructure for neural operator architectures; Optimize training pipelines for throughput, fault tolerance, and cost efficiency; Build serving infrastructure for pre-trained LPMs supporting zero-shot inference and uncertainty quantification; Solve data loading bottlenecks for large-scale mesh datasets; Improve developer experience with reliable CI/CD and debugging tools.
Seniority
Principal, hands-on IC with architectural autonomy