Sr. Staff Software Development Engineer - AI Platform
Core
Design, scale, and maintain enterprise cloud infrastructure and platform capabilities to support production AI/ML workloads, driving infrastructure automation, observability, and platform governance.
Role type
Sr. Staff Platform Engineer (AI Infrastructure)
Builds
Scalable, secure, and highly available AWS infrastructure (EKS, Lambda, ECS, VPC, IAM) and CI/CD pipelines for AI/ML services.
Domain
Cloud Infrastructure / AI Platform Engineering
Deliverable
infrastructure
Required skills
AWS (EKS, Lambda, ECS, VPC, IAM, S3), Terraform, GitLab CI/CD, Prometheus, Grafana, Kubernetes (EKS), Python, Go, Bash, DORA metrics, Incident response, Helm charts, Autoscaling (HPA, Cluster Autoscaler, Karpenter)
Preferred skills
LiteLLM, Ray.io/Anyscale, Temporal, LangGraph, LangChain, LlamaIndex, Vector databases (Qdrant, Pinecone, Weaviate), ZEP, MLflow, Kubeflow, Arize Phoenix, ArgoCD/Flux, Loki, ELK/OpenSearch, Jaeger/Tempo, FinOps, SOC 2, ISO 27001
Responsibilities
Design and maintain scalable AWS infrastructure for AI/ML workloads using Terraform; Own and evolve GitLab CI/CD pipelines for AI platform services; Architect a centralized observability stack using Prometheus and Grafana; Define and track DORA metrics to drive delivery improvements; Implement infrastructure security best practices and platform governance; Mentor junior/mid-level engineers through design and code reviews.
Seniority
Sr. Staff, hands-on IC with strategic ownership
