Senior DevOps Engineer/Site Reliability Engineer-East Coast
Core
Building, operating, and scaling reliable cloud-native infrastructure and distributed data platforms for mission-critical systems.
Role type
Senior hands-on IC DevOps/Site Reliability Engineer
Builds
Cloud-native infrastructure, distributed data platforms, and automation tooling
Domain
Cybersecurity, Cloud Infrastructure, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, Helm, Python, Go, Bash, Linux system administration, CI/CD pipelines, observability (Prometheus, Grafana, Loki, Elastic Stack), incident management, GitOps (ArgoCD, GitHub Actions), distributed data platforms (Kafka, Spark, Elasticsearch, Redis, MongoDB)
Preferred skills
AI-assisted operational tooling, auto-remediation, alert correlation
Technologies
OCI, AWS, GCP, Azure, Docker, Prometheus, Grafana, Loki, Alertmanager, Elastic Stack, ArgoCD, GitHub Actions, Kafka, Spark, Elasticsearch, Redis, MongoDB
Responsibilities
Administer and maintain Kubernetes clusters and containerized workloads; Manage cloud infrastructure across OCI, AWS, GCP, or Azure; Develop and maintain CI/CD pipelines; Implement and manage Infrastructure as Code using Terraform and Helm; Build automation tooling and operational workflows; Drive observability initiatives including monitoring, logging, tracing, and alerting; Monitor, troubleshoot, and resolve production incidents; Support and optimize distributed data platforms; Improve platform reliability, scalability, and operational efficiency using SRE best practices; Perform Linux system administration and networking troubleshooting; Contribute to incident response processes, postmortems, and reliability improvements; Evaluate and implement AI-assisted operational tooling.
Seniority
Senior, hands-on IC
