Software Engineer - SRE
Core
Design, build, and optimize cloud and data infrastructure to ensure high availability, reliability, and scalability of SaaS systems operating at multi-region scale.
Role type
Senior IC Site Reliability Engineer (SRE)
Builds
Cloud platforms, data infrastructure, and automation tools for enterprise customers
Domain
Cloud infrastructure, SaaS, and enterprise data platforms
Deliverable
production ML models | product features | infrastructure
Required skills
Linux systems administration, Public Cloud (AWS), Docker, Kubernetes, Ansible, Networking, Security, Software Development, CI/CD, Infrastructure as Code (Terraform, EKS), Observability (Splunk, Prometheus, Grafana, Thanos, CloudWatch, OpenTelemetry, ELK), Python, Go
Preferred skills
Cloud-based data platform architecture, Multi-priority management in fast-paced environments, High-quality code in Python/Go, Software architecture at scale, Emerging technology research
Technologies
AWS, Docker, Kubernetes, Ansible, Terraform, EKS, Splunk, Prometheus, Grafana, Thanos, CloudWatch, OpenTelemetry, ELK, Python, Go
Responsibilities
Design and optimize cloud infrastructure for high availability; Collaborate with cross-functional teams to create scalable solutions; Troubleshoot production issues and perform root cause analysis; Lead architectural vision and technical strategy; Mentor teams and foster engineering excellence; Engage with customers to translate use cases into actionable insights; Develop strategic roadmaps and infrastructure plans for enterprise deployment
Seniority
Senior, hands-on IC with technical leadership