Senior Site Reliability Engineer
Core
Design, build, and maintain scalable infrastructure to support real-time analytics and machine learning workloads, ensuring reliability and operational excellence.
Role type
Senior Site Reliability Engineer (Infrastructure & ML Workloads)
Builds
Scalable infrastructure for data pipelines, ML workloads, and real-time analytics systems
Domain
Cloud-native infrastructure, Data Engineering, Machine Learning Operations
Deliverable
infrastructure
Required skills
Linux systems administration, Networking, Containerization (Docker), Orchestration (Kubernetes), Infrastructure-as-Code (Terraform, Ansible), Scripting (Python, Bash), Monitoring and Observability (Prometheus, Grafana, Datadog, ELK, OpenTelemetry), CI/CD pipelines
Preferred skills
Cloud managed services (AWS), Data-intensive platforms (Spark, Airflow, Kafka), Security practices for cloud-native applications, High-compliance environments (SOC-2)
Responsibilities
Design and maintain scalable infrastructure for real-time analytics and ML; Improve system reliability through automation and capacity planning; Own and evolve CI/CD pipelines and config management; Implement monitoring, alerting, and incident response processes; Collaborate with engineering and data science teams; Ensure security and compliance; Drive post-incident analysis and continuous improvement
Seniority
Senior, hands-on IC