CareerPlanGet AI match score →

Staff + Sr. Software Engineer, AI Reliability

New York City, NY💼 Full-time💰 $325,000–$325,000🗓 2026-05-05 → 2026-07-31

Core

Design and implement monitoring, observability, and high-availability serving infrastructure for large language model systems to ensure robustness and rapid incident recovery.

Role type

Staff/Senior IC SRE (LLM serving infrastructure)

Builds

Production-grade serving infrastructure and observability systems for Anthropic's AI models

Domain

AI/ML infrastructure, distributed systems, cloud computing

Deliverable

infrastructure

Required skills

distributed systems, infrastructure, reliability engineering, incident response, high-availability architecture, multi-region deployment, cloud provider expertise

Preferred skills

SRE/Production Engineer experience, large-scale model serving (>1000 GPUs), ML hardware accelerators (GPUs, TPUs, Trainium), ML networking (RDMA, InfiniBand), AI observability tools, chaos engineering, open-source infrastructure contributions

Technologies

RDMA, InfiniBand, GPUs, TPUs, Trainium

Responsibilities

Develop Service Level Objectives for LLM serving systems, design monitoring and observability systems across the token path, assist in implementing high-availability serving infrastructure, lead incident response for critical AI services, support reliability of safeguard model serving

Seniority

Staff/Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗