Software Development Engineer
Core
Build and operate large-scale SaaS products like Adobe Document Cloud, ensuring high uptime, service quality, and reliability through operational excellence and AI-driven automation.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Globally distributed multi-cloud environments (AWS, Azure) and AI-enabled reliability automation agents
Domain
Cloud Infrastructure / SaaS / Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (EKS, AKS, ECK), Python, Java, Linux, Cloud automation, Chaos engineering, SLO/SLI definition, Root cause analysis, AI/ML for observability, AIOps, Prompt engineering
Preferred skills
Generative AI agents, GPU infrastructure management, Model serving monitoring, FedRAMP Moderate environment experience
Technologies
Kubernetes, AWS, Azure, Jenkins, Git, Jira, Confluence, Prometheus, Grafana, NewRelic, Splunk, Python, Java, Ruby, MySQL, Tomcat, Memcached, Qpid
Responsibilities
Define and measure service level objectives (SLOs) and indicators (SLIs), Embed with product teams to ensure operational partnership, Automate repeatable tasks at scale, Solve performance and stability issues, Participate in on-call rotation, Mentor junior engineers, Apply AI/ML techniques for anomaly detection and automated remediation
Seniority
Senior, hands-on IC
