Staff Engineer - Site Reliability
Core
Ensure reliability, scalability, and performance of middleware and cloud infrastructure services for an AI-native platform.
Role type
Staff Site Reliability Engineer (Middleware & Cloud Infrastructure)
Builds
Middleware and cloud infrastructure platforms supporting Nextiva's AI-native platform
Domain
Cloud Infrastructure, Distributed Systems, AI/ML Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kafka, Vector Databases, Kubernetes, GCP, Observability, Infrastructure as Code, Automation scripting, Linux administration, Database operations
Preferred skills
AI Native platforms, LLM infrastructure, Vector search, Embedding technologies, Technical leadership
Technologies
Kafka, Vector Databases (Weaviate), GCP, GKE, Terraform, Python, Go, Shell, Datadog, Splunk, OpenTelemetry, Prometheus, Grafana, MongoDB, PostgreSQL, Redis, Elasticsearch, ClickHouse
Responsibilities
Own reliability, availability, scalability, and performance of middleware and cloud infrastructure platforms; Support and optimize Kafka environments; Administer and optimize Vector Database platforms; Manage and support GCP and GKE environments; Drive infrastructure automation and operational excellence; Lead production incident response and root cause analysis; Build and maintain monitoring and observability solutions
Seniority
Staff, hands-on IC with strategic impact
