System Engineer (Cloud & AI)
Core
Design, deploy, and operate multi-cloud and on-prem infrastructure while integrating AI tooling into operations workflows.
Role type
Senior System Engineer (Cloud & AI)
Builds
Production-ready cloud infrastructure, automated operations pipelines, and AI-augmented runbooks.
Domain
Cloud Infrastructure & AI Operations
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, Infrastructure as Code (Terraform), multi-cloud management (AWS, GCP), container orchestration (Kubernetes), scripting (Python, Bash, Go), observability (Prometheus, Grafana), security hardening
Preferred skills
AI/ML workload support, self-service automation design, cost optimization strategies
Technologies
Terraform, AWS (EC2, RDS, S3, EKS, ECS), GCP (Compute Engine, Cloud SQL, GKE), Docker, Kubernetes, Prometheus, Grafana, Python, Bash, Go
Responsibilities
Own reliability of multi-environment platform across AWS, GCP, and on-prem; Build and operate infrastructure as code; Improve cost, scalability, and performance; Operate and harden production systems; Run critical internet-facing foundations; Automate operations with scripts/playbooks; Improve observability and reliability
Seniority
Senior, hands-on IC