Principal Engineer
Core
Lead the development of critical interfaces for managing system state, bridging hardware, AI software, and cloud environments to enable reliable large-scale AI deployment.
Role type
Principal Engineer (System Management)
Builds
Interfaces between hardware, AI software, and frameworks; system management capabilities for public and private cloud environments.
Domain
AI Infrastructure / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Linux-based infrastructure design, distributed systems, technical leadership, architecture decision-making, RESTful API design, Go, Bash, Python, Kubernetes, Infrastructure-as-Code, hardware management interfaces (Redfish/IPMI), Linux systems engineering, technical mentoring.
Preferred skills
AI coding assistants, Kubernetes operators, HPC environments (SLURM/LSF), virtualization (Open vSwitch/KVM/QEMU), distributed storage (Ceph), observability (Grafana/Prometheus/OpenTelemetry), managed network switches (EOS/SONiC/DNOS), AI infrastructure/PyTorch support.
Technologies
Go, Bash, Python, Kubernetes, Terraform/OpenTofu, Ansible, GitLab, GitHub Actions, Git, Redfish, IPMI, SLURM, LSF, Open vSwitch, KVM, QEMU, Ceph, Grafana, Prometheus, OpenSearch, Loki, Mimir, OpenTelemetry, EOS, SONiC, DNOS, PyTorch.
Responsibilities
Convert system management direction into technical plans and priorities; provide technical leadership for architecture and design decisions; act as a technical authority for complex engineering initiatives; provide technical direction and mentoring to engineers; take responsibility for key technical outcomes across the full software lifecycle; identify and lead improvements for reliability, scalability, and operability; collaborate with cross-functional teams to diagnose system-level issues; improve engineering standards including CI/CD and automation; act as a senior technical escalation point.
Seniority
Principal, hands-on IC with strategic leadership