Sr. Applied Scientist, Infrastructure Reliability
Core
Architect and build agentic observability solutions that autonomously detect anomalies and drive automated remediation for Amazon's critical infrastructure.
Role type
Senior Applied Scientist (Infrastructure Reliability)
Builds
Production ML systems for network telemetry analysis, anomaly detection, and automated incident response.
Domain
Cloud Infrastructure / Observability / AI
Deliverable
production ML models
Required skills
Machine learning, deep learning, time-series analysis, graph-based methods, agentic AI architectures, multi-agent orchestration, LLMs, root cause analysis, predictive failure detection
Preferred skills
Anomaly detection, natural language processing, information retrieval, modeling tools (R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy)
Technologies
Java, C++, Python, Tensorflow, numpy, scipy, Spark MLLib, MxNet, R, scikit-learn
Responsibilities
Define science vision for agentic observability and translate it into research and engineering roadmaps; Build ML systems that autonomously detect, classify, and correlate infrastructure anomalies; Design models for event correlation and predictive failure detection; Own agentic architecture for automated observability workflows; Curate datasets and evaluation methodologies for detection systems; Write production-quality code for critical-path components; Mentor scientists and engineers.
Seniority
Senior, hands-on IC with research leadership