CareerPlanGet AI match score →

Site Reliability Engineer (SRE)

San Francisco💼 Full-time💰 $350,000–$350,000🗓 2026-05-04 → 2026-07-31

Core

Drive end-to-end reliability for Tinker, a fine-tuning API for customizing frontier AI models, ensuring robust distributed training systems and multi-tenant isolation.

Role type

Site Reliability Engineer (SRE)

Builds

Tinker platform (fine-tuning API for open weights models)

Domain

AI/ML Infrastructure, Distributed Systems

Deliverable

production ML models

Required skills

distributed systems, cloud infrastructure, site reliability engineering, software development for tooling/automation, production incident response, postmortems, systematic reliability improvement

Preferred skills

operating production cloud services at scale, distributed training frameworks, checkpoint and recovery systems for long-running jobs, Kubernetes at scale with heterogeneous GPU workloads

Technologies

Kubernetes, public cloud platforms, distributed training frameworks

Responsibilities

Define and own end-to-end reliability from CI/CD to observability; Develop Service Level Objectives for distributed training systems; Design and implement monitoring across the full training path; Drive incident response and postmortems; Harden multi-tenant isolation and resource scheduling; Collaborate with security teams on production vulnerabilities

Seniority

Mid-to-Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗