Senior Software Engineer, AIOps
Core
Building a mission-critical Observability and Prediction platform (SaaS and on-premises) that ingests massive telemetry streams from GPU clusters and operationalizes predictive AI models at scale.
Role type
Senior Software Engineer (Distributed Systems & Production ML)
Builds
High-scale distributed systems for telemetry ingestion, agentic AIOps monitoring, and model-serving infrastructure.
Domain
AI Infrastructure / GPU Clusters / Observability
Deliverable
production ML models | product features | infrastructure
Required skills
Go, C++, Rust, Kubernetes, distributed systems design, high-performance concurrent architectures, ML model deployment, data storage strategies
Preferred skills
ML model-serving platforms, MLOps tooling, prototype-to-production track record, full-stack systems thinking
Technologies
Go, C++, Rust, Kubernetes
Responsibilities
Architect agentic AIOps systems for GPU fleet health monitoring; design distributed systems for extreme telemetry density; instrument services with deep observability; build and own model-serving infrastructure; contribute to core platform libraries.
Seniority
Senior, hands-on IC