Site Reliability Engineer (SRE) Intern — AI Infrastructure
Core
Support daily operations of AI infrastructure including servers, storage, networking, and software components.
Role type
SRE Intern (AI Infrastructure)
Builds
High-performance AI compute environments (bare-metal, Kubernetes, Slurm)
Domain
AI Infrastructure / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Linux system administration, server hardware diagnostics, Bash/Python scripting, configuration management, CI/CD tools, observability tools (Prometheus, Grafana, ELK), cluster software (Slurm, Kubernetes)
Preferred skills
AI/ML infrastructure experience, distributed systems knowledge, networking fundamentals
Responsibilities
Deploy and maintain AI infrastructure servers and networking equipment, perform hardware diagnostics and firmware updates, troubleshoot onsite operational issues, document incidents and resolutions
Seniority
Intern