CareerPlanSign in

Machine Learning Infrastructure Engineer

Redwood City, CA💼 Full-time🗓 2025-04-24 → 2026-09-25

Core

Designing, building, and maintaining training and serving infrastructure for ML research.

Role type

ML Infrastructure Engineer

Builds

Training and serving infrastructure for ML research

Domain

Cloud infrastructure, High-performance computing, GPU clusters

Deliverable

infrastructure

Required skills

Cloud platform management, Kubernetes, GPU cluster management, Tooling development for diagnostics, Experiment management, GPU utilization optimization

Preferred skills

Large GPU cluster management, High-performance networking, LLM training support, ML framework development (PyTorch/TensorFlow/JAX), GPU kernel development

Technologies

Compute Engine, Kubernetes, Cloud Storage, PyTorch, TensorFlow, JAX

Responsibilities

Provide infrastructure support to ML research and product teams, Build tooling to diagnose cluster issues and hardware failures, Monitor deployments and manage experiments, Maximize GPU allocation and utilization for serving and training

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.