Service Operations Engineer
Core
Ensure 24/7 stability of products and infrastructure by monitoring services, responding to incidents, and analyzing root causes.
Role type
Service Operations Engineer (SRE)
Builds
Stable internal and customer-facing tech products for a venture builder
Domain
Technology Operations / Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux administration, Command Line Interface (CLI), TCP/IP networking, HTTP, SQL (SELECT queries), Incident troubleshooting, Root cause analysis
Preferred skills
Bash scripting, Python scripting, YouTrack, Jira, Prometheus, Grafana, DataDog, Technical support experience
Technologies
Linux, Bash, Python, YouTrack, Jira, Prometheus, Grafana, DataDog, SQL
Responsibilities
Monitor products and infrastructure 24/7, Detect and escalate user-impacting incidents, Collaborate with development and infrastructure teams, Maintain event logs and incident documentation, Develop and maintain monitoring/alerting systems, Conduct technical investigations and root cause analysis
Seniority
Individual Contributor