Sr Staff Site Reliability Engineer
Core
Design, implement, and maintain robust infrastructure and automation solutions for an internal LLM-powered chat service and core systems.
Role type
Sr Staff Site Reliability Engineer
Builds
Cloud-native infrastructure, data pipelines, CI/CD pipelines, and observability strategies for an aerospace company's internal services.
Domain
Aerospace / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Amazon EKS, observability (Prometheus, Grafana, ELK), data pipelines (Kafka, Airflow, Spark), CI/CD (Jenkins, GitLab CI, ArgoCD), Docker, Python, Bash, PowerShell, distributed systems, networking
Preferred skills
Other Kubernetes distributions, compliance frameworks (SOC 2, HIPAA, GDPR), AWS certifications
Technologies
Amazon EKS, OpenRouter, Prometheus, Grafana, ELK, Kafka, Airflow, Spark, Jenkins, GitLab CI, ArgoCD, Docker, Python, Bash, PowerShell
Responsibilities
Implement and maintain infrastructure for an internal LLM-powered chat service; manage highly available, scalable, and secure cloud-native infrastructure on EKS; develop observability strategies; architect and optimize data pipelines; drive CI/CD improvements; enforce security practices; design Docker-based containerization; develop automation scripts; collaborate on reliability in SDLC; troubleshoot production issues; participate in on-call rotations
Seniority
Sr Staff, hands-on IC with mentorship