Sr. Site Reliability Engineer
Core
Guardian of production ecosystems for complex, data-driven AI platforms, ensuring resilience, scalability, and high performance through a hybrid of software engineering and systems architecture.
Role type
Senior Site Reliability Engineer (MLOps focus)
Builds
Production-grade AI/ML services, scalable Kubernetes clusters, and resilient ML pipelines
Domain
Artificial Intelligence / Machine Learning Operations / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (K8s), Docker, Terraform, Python, Bash, CI/CD, Observability, Incident Response, GPU/TPU optimization
Preferred skills
Go, Kubeflow, Vertex AI, MLflow, DVC, Vector Databases, Istio/Anthos, BigQuery, Pub/Sub
Technologies
Kubernetes, GKE, Vertex AI, Kubeflow, MLflow, DVC, Terraform, Pulumi, GitHub Actions, Cloud Build, ArgoCD, Prometheus, Grafana, Stackdriver, BigQuery, Pub/Sub, Pinecone, Milvus, Istio, Anthos
Responsibilities
Define and maintain SLOs/SLIs for critical AI/ML services; Architect and manage auto-scaling strategies for Kubernetes; Ensure high availability of model serving endpoints; Optimize GPU/TPU resource utilization; Stabilize ML pipelines; Provision and manage cloud environments via IaC; Design and optimize deployment pipelines; Develop automation scripts for operational tasks; Build and manage observability dashboards; Lead incident response and root cause analysis
Seniority
Senior, hands-on IC
