MTS - Site Reliability Engineer
Core
Ensure uptime, resiliency, and fault tolerance of AI model training and inference systems while building automation for hybrid cloud/on-prem environments.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
Production-ready AI model training and inference pipelines on hybrid cloud/on-prem CPU+GPU infrastructure
Domain
Generative AI, Cloud Infrastructure, High-Performance Computing
Deliverable
production ML models
Required skills
Kubernetes, Docker, CI/CD pipelines, public cloud platforms (Azure/AWS/GCP), infrastructure-as-code, monitoring & observability tools (Grafana, Datadog, OpenTelemetry), Python/Go/Bash scripting, distributed systems, networking, storage
Preferred skills
Large-scale GPU cluster management, ML training/inference pipelines, HPC workload schedulers, capacity planning, cost optimization for GPU-heavy environments
Technologies
Kubernetes, Docker, Azure, AWS, GCP, Grafana, Datadog, OpenTelemetry
Responsibilities
Design and maintain monitoring, alerting, and logging systems for model serving pipelines; Analyze system performance and optimize resource utilization (compute, GPU clusters, storage, networking); Build automation for deployments, incident response, scaling, and failover; Lead on-call rotations and conduct blameless postmortems; Partner with ML engineers to accelerate research-to-production workflows
Seniority
Senior, hands-on IC