Sr Software Engineer (Site Reliability) Austin or Dallas, TX
Core
Design, build, and maintain scalable, reliable distributed systems and cloud infrastructure for the Digital Fulfillment team, ensuring high availability and performance.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Cloud infrastructure, CI/CD pipelines, monitoring tooling, and microservices platforms
Domain
Cloud Infrastructure & Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Distributed systems design, Kubernetes (GKE/K8s), Terraform, CI/CD pipelines, Java (Spring), PostgreSQL, Linux, REST/GraphQL APIs, Python/Ruby/Bash scripting, Scalability & High Availability principles
Preferred skills
Google Kubernetes Engine (GKE), AWS environments, GitLab/GitHub Actions, Datadog/Grafana/New Relic, Microservices architecture patterns
Technologies
Kubernetes, Terraform, GCP, AWS, GitLab, GitHub Actions, PostgreSQL, Docker, Linux, Datadog, Grafana, New Relic, Java, Python, Ruby, Bash, Groovy
Responsibilities
Develop and maintain tooling for environment monitoring and task automation; Analyze and establish efficient configurations for software, servers, and databases; Collaborate with development teams on service architecture and capacity planning; Monitor SLOs/SLAs and resolve gaps; Troubleshoot production support and on-call issues; Serve as technical SME for cross-functional engineering teams
Seniority
Senior, hands-on IC