Infrastructure Software Engineer
Core
Lead the development of next-generation infrastructure tooling, hybrid HPC clusters, and observability stacks to support AI ASIC development, simulation, and CI workflows.
Role type
Senior Infrastructure Software Engineer (Systems & Observability)
Builds
Hybrid HPC clusters, programmable infrastructure control planes, real-time telemetry/observability systems, and synthetic testing frameworks.
Domain
AI Infrastructure / Semiconductor (ASIC) / High-Performance Computing
Deliverable
infrastructure
Required skills
Systems programming, Linux administration, containerization, CI/CD pipeline design, Infrastructure as Code, observability architecture, distributed systems debugging, cloud/hybrid deployment strategies
Preferred skills
ASIC development flows (Synopsys, Cadence, Verilator), Bazel build system, cloud provider expertise (AWS/GCP/Azure), bare-metal server management, high-performance storage systems, telemetry system optimization
Technologies
Python, Go, Rust, C++, Terraform, Ansible, Puppet, OpenTofu, SLURM, Kubernetes, Prometheus, Grafana, VictoriaMetrics, Loki, Jupyter, VS Code
Responsibilities
Design and build orchestration layers for hybrid HPC clusters; develop a programmable infrastructure control plane; create tools for massive parallelism; prototype workload migration strategies between on-prem and cloud; implement real-time telemetry and tracing systems; build a full observability stack with synthetic testing.
Seniority
Senior, hands-on IC with mentorship responsibilities