Cloud Platform Engineer
Core
Guardian of reliability, performance, and scalability for the AI Inferencing Service, bridging software development and operations to ensure exceptional uptime and low-latency response times.
Role type
Senior Cloud Platform Engineer (AI Infrastructure)
Builds
Production AI inferencing endpoints and supporting infrastructure for enterprise/government generative AI platforms
Domain
High-performance computing, Generative AI, Cloud Infrastructure
Deliverable
production ML models
Required skills
Python, Go, Java, Kubernetes, Docker, Terraform, Ansible, Prometheus, Grafana, Datadog, CI/CD pipeline design, SLO/SLI definition, capacity planning, incident management, Linux system administration
Preferred skills
Hybrid cloud/on-prem infrastructure, ML/AI inferencing support, GPU-accelerated computing (NVIDIA), model serving frameworks (vLLM, SGLang, Ray), MLOps, database management (SQL/NoSQL), caching systems (Redis, Memcached)
Technologies
AWS, GCP, Azure, Prometheus, Grafana, Datadog, Terraform, Ansible, Jenkins, GitHub Actions, ArgoCD, Docker, Kubernetes, vLLM, SGLang, Ray, Redis, Memcached
Responsibilities
Manage production inferencing service availability, latency, and performance across multiple regions; participate in shared on-call rotation for 24/7 support; develop and maintain monitoring, alerting, and dashboarding; design and implement auto-scaling policies; manage cloud infrastructure using IaC; build and improve CI/CD pipelines; forecast infrastructure needs and optimize cloud costs; define and report on SLOs and SLIs.
Seniority
Senior, hands-on IC