Member of the Technical Staff - Systems ML Engineer
Core
Optimize large-scale ML model training and inference performance across cloud and edge infrastructure for physical AI systems in scientific research.
Role type
Senior IC Systems ML Engineer
Builds
High-performance ML training pipelines, custom GPU kernels, and production inference systems for scientific data
Domain
Frontier science, machine learning, robotics, and cloud infrastructure
Deliverable
production ML models
Required skills
Large-scale model optimization, custom GPU kernel development (CUDA/Triton), distributed training frameworks (PyTorch/JAX), cloud infrastructure management (AWS, Kubernetes), data pipeline engineering, performance profiling (Nsight, PyTorch Profiler), cost optimization
Preferred skills
Edge device deployment, security and compliance implementation, monitoring and versioning systems
Technologies
PyTorch, JAX, CUDA, Triton, AWS, Terraform, Kubernetes, Nsight
Responsibilities
Profile and optimize training/inference workloads, develop custom GPU kernels, manage cloud and edge infrastructure, debug complex stack failures, build data pipelines, implement deployment safety mechanisms
Seniority
Senior, hands-on IC
