Platform & SRE Engineer
Core
Own the operational lifecycle of Linux-based runtime hosts, manage AWS infrastructure, and establish SRE practices for a distributed AI runtime platform.
Role type
Senior Platform & SRE Engineer (via careerplan.io/jobs/61-005-79-371-platform-sre-engineer-at-mwdn)
Builds
High-performance, low-latency distributed runtime environments for production AI systems
Domain
AI Infrastructure / Real-Time Data Processing
Deliverable
infrastructure
Required skills
Linux systems administration, AWS EC2 and AMI management, Linux security hardening, observability stack (Grafana, Prometheus, OpenTelemetry), SRE principles, Python and Bash automation, networking fundamentals
Preferred skills
Linux performance tuning (CPU pinning, NUMA), DPDK, FPGA integration, Infrastructure as Code (Terraform), CI/CD (GitHub Actions), vulnerability management
Technologies
Linux, systemd, AWS, EC2, AMI, Grafana, OpenTelemetry, Prometheus, Mimir, ClickHouse, Fluent Bit, Python, Bash, TCP/IP, TLS
Responsibilities
Own the operational lifecycle of Linux-based runtime hosts, Build and maintain production dashboards and alerts, Define and implement SLIs and SLOs, Automate host provisioning and image lifecycle management, Harden operating-system configuration and services, Troubleshoot complex issues across OS, networking, and hardware
