Senior Production Engineer, Core PE
Core
Ensure reliability, scalability, and performance of Crusoe's GPU cloud platform powering next-generation AI workloads.
Role type
Senior Production Engineer (Operational Excellence)
Builds
Reliable, energy-efficient, AI-optimized cloud platform
Domain
AI infrastructure / GPU cloud / HPC
Deliverable
production ML models
Required skills
Linux/Unix debugging, Kubernetes, distributed systems, observability (Prometheus, Grafana, OpenTelemetry), infrastructure-as-code (Terraform, Ansible), scripting (Go, Python, C, C++)
Preferred skills
GPU workload support, self-healing system design, incident management frameworks (SRE, ITIL)
Technologies
Prometheus, Grafana, Alertmanager, OpenTelemetry, Kubernetes, Terraform, Ansible, AWS, GCP
Responsibilities
Define and evolve availability metrics (SLIs/SLOs), respond to production incidents and perform root cause analysis, build observability tooling, identify reliability risks and bottlenecks, develop automation to reduce operational toil, partner with compute/networking/storage teams on resilience
Seniority
Senior, hands-on IC