Senior Site Reliability Engineer
Core
Keep production systems running smoothly by fusing engineering principles, operational knowledge, security, and automation to ensure platform/service production excellence for AI and ML workloads.
Role type
Senior Site Reliability Engineer (Infrastructure & AI Platform)
Builds
AI Platform Core services, infrastructure for public cloud deployment, and software delivery lifecycle tooling for ML/LLM features.
Domain
Cloud Infrastructure, DevOps, Machine Learning Operations (MLOps), Large Language Model Operations (LLMOps)
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Infrastructure as Code (Terraform/CloudFormation), Python or Go, Observability (metrics/logging/tracing), Incident response, Cost governance, Security compliance
Preferred skills
GPU scheduling, Model serving (KServe/RayServe/Triton/vLLM), MLOps pipelines (Kubeflow/MLflow/Feast), LLMOps (prompt management/evals/guardrails), Cross-functional collaboration
Technologies
Kubernetes, Helm, Terraform, AWS, Python, Go, Grafana, Istio, GitHub Actions, GitOps, KServe, RayServe, Triton, vLLM, Kubeflow, MLflow, Feast, W&B
Responsibilities
Manage reliability of ML/LLM workloads including model serving, inference infrastructure, autoscaling, and SLOs; Implement observability for ML models including drift monitoring, tracing, evals, and guardrails; Shape company-wide technical direction and build reusable developer tooling and automation; Package reusable components for open-source tools and ML infrastructure; Enforce secure-by-default infrastructure with compliance audits and cost governance.
Seniority
Senior, hands-on IC