Site Reliability Engineer
Core
Maintain and operate Kubernetes clusters and cloud-based infrastructure to strengthen production readiness and reliability for a large-scale software platform.
Role type
Site Reliability Engineer (SRE)
Builds
Cloud infrastructure and monitoring solutions
Domain
Automotive and Smart Mobility
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Infrastructure as Code (IaC), public cloud platforms, incident response, system troubleshooting, automation
Preferred skills
Go, Python, OpenTelemetry, Prometheus, APM tools, disaster recovery planning, chaos engineering, capacity management
Responsibilities
Maintain and operate Kubernetes clusters and cloud-based infrastructure; Design and build software solutions that improve monitoring and service reliability; Drive productivity by automating workflows; Participate in on-call rotations to monitor system health and respond to incidents; Provide technical support and resolution for escalated production issues
Seniority
Mid-level, hands-on IC