CareerPlanSign in

Software Engineer, SRE

USA💼 Full-time💰 $211,000–$211,000🗓 2026-09-09 → 2026-10-03

Core

Designing and implementing scalable, reliable, and secure AWS cloud infrastructure and observability stacks for complex SaaS systems and LLM deployments.

Role type

Senior Site Reliability Engineer (Infrastructure & Observability)

Builds

Production-grade AWS infrastructure, observability platforms, and CI/CD pipelines for LLM and SaaS applications.

Domain

Cloud Infrastructure, SaaS, Large Language Models (LLM)

Deliverable

production ML models | infrastructure

Required skills

AWS services, Terraform, container orchestration, cloud networking (IAM, VPC), observability (Prometheus, Grafana, Datadog), CI/CD pipeline design, incident management, system scalability design.

Preferred skills

Enterprise customer compliance knowledge, cross-functional collaboration with product and ML teams, SRE culture definition.

Technologies

AWS, Terraform, Prometheus, Grafana, Datadog, CI/CD tools, LLM frameworks

Responsibilities

Own the observability stack (monitoring, alerting, logging, tracing); design reliable and scalable systems with product/platform engineers; implement secure AWS infrastructure using Terraform; improve reliability and scalability of LLM deployments; lead enhancements to deployment pipelines and incident management processes; define SRE practices and best practices across engineering.

Seniority

Senior, hands-on IC with strategic influence

Sourced via devitjobs · Listed on CareerPlan, which tracks 937,000+ jobs from 20+ sources.