Software Engineer, SRE
Core
Designing and implementing scalable, reliable, and secure AWS cloud infrastructure and observability stacks for complex SaaS systems and LLM deployments.
Role type
Senior Site Reliability Engineer (Infrastructure & Observability)
Builds
Production-grade AWS infrastructure, observability platforms, and CI/CD pipelines for LLM and SaaS applications.
Domain
Cloud Infrastructure, SaaS, Large Language Models (LLM)
Deliverable
production ML models | infrastructure
Required skills
AWS services, Terraform, container orchestration, cloud networking (IAM, VPC), observability (Prometheus, Grafana, Datadog), CI/CD pipeline design, incident management, system scalability design.
Preferred skills
Enterprise customer compliance knowledge, cross-functional collaboration with product and ML teams, SRE culture definition.
Technologies
AWS, Terraform, Prometheus, Grafana, Datadog, CI/CD tools, LLM frameworks
Responsibilities
Own the observability stack (monitoring, alerting, logging, tracing); design reliable and scalable systems with product/platform engineers; implement secure AWS infrastructure using Terraform; improve reliability and scalability of LLM deployments; lead enhancements to deployment pipelines and incident management processes; define SRE practices and best practices across engineering.
Seniority
Senior, hands-on IC with strategic influence