Senior Site Reliability Engineer
Core
Design, build, and maintain scalable infrastructure supporting real-time analytics and machine learning workloads, ensuring reliability and operational excellence.
Role type
Senior Site Reliability Engineer (Infrastructure & ML Platforms)
Builds
Scalable infrastructure, data pipelines, ML workloads, and real-time analytics systems
Domain
Cloud Infrastructure, Machine Learning, Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, containerization (Docker), orchestration (Kubernetes), infrastructure-as-code (Terraform, Ansible), CI/CD pipeline management, monitoring and observability (Prometheus, Grafana, Datadog, ELK, OpenTelemetry), scripting (Bash, Python), incident response, capacity planning
Preferred skills
Cloud managed services (AWS), data-intensive platforms (Spark, Airflow, Kafka), security practices for cloud-native applications, high-compliance environments (SOC-2)
Responsibilities
Design and maintain scalable infrastructure for ML and analytics; improve system reliability via automation and observability; own and evolve CI/CD pipelines; implement monitoring, alerting, and incident response processes; collaborate with engineering and data science teams; ensure security and compliance; drive post-incident analysis
Seniority
Senior, hands-on IC