Staff Software Engineer, Inference Platform
Core
Lead the orchestration layer for a globally distributed inference platform running on Cerebras AI chips, connecting cloud components with machine learning services.
Role type
Staff Software Engineer (Inference Platform)
Builds
Production software for a globally distributed inference platform, Kubernetes operators, and service security policies.
Domain
AI Infrastructure / Distributed Systems / Cloud Infrastructure
Deliverable
production ML models
Required skills
Distributed systems architecture, Kubernetes, Backend or systems languages (Go or C++), Security (TLS, mTLS), Observability (metrics, logging, tracing, alerting), Incident response, SLO-driven operations, High-consequence architectural decision making
Preferred skills
ML inference infrastructure, Model serving systems, GPU-accelerated workloads, TTFT and tail-latency reduction optimization
Technologies
Kubernetes, Go, C++, TLS, mMetrics, Logging, Tracing
Responsibilities
Design, develop, test, and maintain production software spanning testing, continuous development, observability, security, networking, debugging, and productionization. Raise the effectiveness of senior engineers through design feedback, pairing, and clear technical standards. Help shape the technical direction for the Inference Platform, Kubernetes custom resource definitions, failure domains, service boundaries, and system evolution. Architect active-active systems with rapid failover, graceful degradation, and clear SLOs. Write and review production code in the most important parts of the platform. Lead on the hardest production issues and cross-system bottlenecks. Partner with ML, Product, Infrastructure, and Cloud teams to translate product and business requirements into scalable system designs.
Seniority
Staff, hands-on IC with strategic influence