Software Engineer, Observability
Core
Building AI-powered observability infrastructure and tools to monitor, debug, and ensure reliability of OpenAI's massive-scale production systems.
Role type
Software Engineer (Observability Infrastructure & AI Tools)
Builds
Distributed logging, time series, and trace storage systems; AI-native agents for SEV summarization and auto-generated dashboards.
Domain
AI Infrastructure / Observability / Cloud Systems
Deliverable
production ML models | infrastructure | product features
Required skills
Large-scale distributed systems, logging systems, time series databases, full-stack development, systems fundamentals, networking, cloud infrastructure
Preferred skills
Observability systems (Prometheus, OpenTelemetry), AI/ML integration, notebook-like UI development
Technologies
Kubernetes, AWS, Prometheus, OpenTelemetry
Responsibilities
Own core observability infrastructure (logging, time series, trace storage); Build AI-native tools for autonomous issue detection and resolution; Contribute to UI experiences (dashboards, interactive debugging); Collaborate with research and product teams to define AI-powered observability features.
Seniority
Mid-to-Senior, hands-on IC