CareerPlanGet AI match score →

Senior Software Engineer, AIOps

2 Locations💼 Full-time🗓 2026-06-15 → 2026-08-01

Core

Building a mission-critical Observability and Prediction platform (SaaS and on-premises) that ingests massive telemetry streams from GPU clusters and operationalizes predictive AI models at scale.

Role type

Senior Software Engineer (Distributed Systems & Production ML)

Builds

High-scale distributed systems for telemetry ingestion, agentic AIOps monitoring, and model-serving infrastructure.

Domain

AI Infrastructure / GPU Clusters / Observability

Deliverable

production ML models | product features | infrastructure

Required skills

Go, C++, Rust, Kubernetes, distributed systems design, high-performance concurrent architectures, ML model deployment, data storage strategies

Preferred skills

ML model-serving platforms, MLOps tooling, prototype-to-production track record, full-stack systems thinking

Technologies

Go, C++, Rust, Kubernetes

Responsibilities

Architect agentic AIOps systems for GPU fleet health monitoring; design distributed systems for extreme telemetry density; instrument services with deep observability; build and own model-serving infrastructure; contribute to core platform libraries.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗