Infrastructure Engineer
Core
Design, build, and maintain scalable cloud infrastructure and platform systems specifically tailored for AI/LLM workloads and agentic systems.
Role type
Senior Infrastructure Engineer (Cloud & AI/LLM)
Builds
Scalable, reliable, and secure cloud infrastructure powering AI/LLM workloads and agentic systems
Domain
Cloud Infrastructure, AI/LLM, DevOps
Deliverable
production ML models | infrastructure
Required skills
AWS services, Terraform, Kubernetes (EKS), CI/CD pipelines, observability, networking fundamentals, scripting (TypeScript, Bash), AI agent frameworks (LangChain, LlamaIndex, CrewAI), LLM deployment
Preferred skills
Golang, GitOps (ArgoCD, Flux), service mesh (Istio, Linkerd), FinOps
Technologies
AWS (VPC, EC2, RDS, S3, IAM, Lambda, EKS), Terraform, Kubernetes, Docker, GitHub Actions, GitLab CI, ArgoCD, Prometheus, Grafana, Datadog, ELK, LangChain, LlamaIndex, CrewAI
Responsibilities
Design and provision cloud infrastructure for AI/LLM workloads; Develop and maintain Terraform modules; Manage Kubernetes clusters and workload orchestration; Build and optimize CI/CD pipelines; Implement logging, monitoring, and alerting solutions; Write scripts to automate operational tasks; Create and maintain runbooks and architecture diagrams; Partner with development teams to improve platform reliability; Participate in on-call rotation for production incidents; Evaluate and integrate AI-driven DevOps tools and LLM-assisted automations
Seniority
Senior, hands-on IC