Infrastructure engineer
Core
Design and operate scalable, fault-tolerant infrastructure across AWS, GCP, and Azure, leading incident response and shaping multi-year platform investments.
Role type
Senior Infrastructure Engineer (Cloud & AI Tooling)
Builds
End-to-end infrastructure builds, observability platforms, and AI-assisted workflows for incident investigation and code review.
Domain
Cloud Infrastructure (AWS, GCP, Azure) and AI-assisted DevOps
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Kubernetes, Helm, Terraform, Prometheus, Grafana, ELK, SLO management, incident response, root-cause analysis, system design
Preferred skills
Pulumi, AI-assisted/agentic tooling development, non-trivial production software design
Technologies
AWS, GCP, Azure, Kubernetes, Helm, Terraform, Pulumi, Prometheus, Grafana, ELK
Responsibilities
Design and operate scalable, fault-tolerant infrastructure; Automate infrastructure management using Python or Go; Lead incident response and post-mortems; Shape multi-year platform investments in observability and reliability; Collaborate on reliable system design; Use and develop AI-assisted workflows for incident investigation and tooling.
Seniority
Senior, hands-on IC