CareerPlanSign in

Member of Technical Staff | ML Systems

Brazil🌐 Remote💼 Full-time🗓 2026-09-23 → 2026-09-25

Core

Build and operate the infrastructure for distributed machine learning training, model serving, and governance, enabling researchers to ship reliable models at scale.

Role type

Senior IC ML Systems Engineer (Distributed Systems & GPU)

Builds

High-performance distributed training infrastructure, experiment tracking systems, model registries, and governed release pipelines.

Domain

Machine Learning Infrastructure / Distributed Systems / GPU Computing

Deliverable

production ML models | infrastructure

Required skills

Distributed machine learning training, CUDA kernel development, GPU performance optimization, columnar data formats (Lance, Arrow), experiment tracking, model lineage and versioning, Python development, multi-node workload management, software engineering principles for production systems

Preferred skills

Graph neural networks, graph sampling, multi-cloud GPU infrastructure (SkyPilot), model governance in regulated environments

Technologies

CUDA, Ray, PyTorch Distributed, Lance, Arrow, CSR/CSC, SkyPilot

Responsibilities

Build high-performance CUDA kernels for graph neural networks; Evolve distributed sampling and training infrastructure; Develop systems for efficient training and serving of ML models at scale; Define data formats and own data materialization pipelines; Establish data contracts with dataset teams; Build experiment tracking and evaluation infrastructure; Own model registry, lineage, and versioning; Define and operate release gates for governed model releases; Ensure full auditability of model releases; Optimize training throughput and GPU utilization; Build internal platform capabilities with a product mindset

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.