Site Reliability Engineer (SRE)
Core
Drive end-to-end reliability for Tinker, a fine-tuning API for customizing frontier AI models, ensuring robust distributed training systems and multi-tenant isolation.
Role type
Site Reliability Engineer (SRE)
Builds
Tinker platform (fine-tuning API for open weights models)
Domain
AI/ML Infrastructure, Distributed Systems
Deliverable
production ML models
Required skills
distributed systems, cloud infrastructure, site reliability engineering, software development for tooling/automation, production incident response, postmortems, systematic reliability improvement
Preferred skills
operating production cloud services at scale, distributed training frameworks, checkpoint and recovery systems for long-running jobs, Kubernetes at scale with heterogeneous GPU workloads
Technologies
Kubernetes, public cloud platforms, distributed training frameworks
Responsibilities
Define and own end-to-end reliability from CI/CD to observability; Develop Service Level Objectives for distributed training systems; Design and implement monitoring across the full training path; Drive incident response and postmortems; Harden multi-tenant isolation and resource scheduling; Collaborate with security teams on production vulnerabilities
Seniority
Mid-to-Senior, hands-on IC