Cloud Platform Engineer
Core
Guardian of reliability, performance, and scalability for a full-stack generative AI inferencing service, bridging software development and operations to ensure exceptional uptime and low-latency response times.
Role type
Senior Cloud Site Reliability Engineer (AI Inferencing)
Builds
Production AI inferencing endpoints and the underlying cloud infrastructure for enterprise and government organizations.
Domain
Generative AI, High-Performance Computing, Cloud Infrastructure
Deliverable
production ML models
Required skills
Site Reliability Engineering, Cloud Infrastructure Management, Container Orchestration, Infrastructure as Code, CI/CD Pipeline Design, Monitoring and Observability, Capacity Planning, Incident Management, Python/Go/Java Programming
Preferred skills
Hybrid cloud/on-prem experience, ML/AI inferencing support, GPU-accelerated computing, Model serving frameworks (vLLM, SGLang, Ray), MLOps, Database and caching management
Technologies
AWS, GCP, Azure, Docker, Kubernetes, Terraform, Ansible, Prometheus, Grafana, Datadog, ELK Stack, Jenkins, GitHub Actions, ArgoCD, Redis, Memcached
Responsibilities
Manage production inferencing service availability, latency, and efficiency across multiple regions; Lead incident response and drive blameless post-mortems; Develop and maintain advanced monitoring and alerting systems; Design and implement auto-scaling policies; Manage cloud infrastructure using IaC; Build and improve CI/CD pipelines for model version deployment; Forecast infrastructure needs and optimize cloud costs.
Seniority
Mid-Senior, hands-on IC (via careerplan.io/jobs/6105132004-cloud-platform-engineer-at-sambanovasystems)
