Senior Software Engineer, Inference Platform
Core
Design and build aion's inference service platform, the backbone for serving AI models at scale across diverse workloads.
Role type
Senior IC backend engineer (inference platform)
Builds
AI Gateway, Resource Orchestrator, Runtime Engines, Autoscaler, and deployment pipelines for model serving
Domain
Enterprise AI / Distributed Systems / GPU Computing
Deliverable
production ML models
Required skills
Golang, distributed systems design, microservices architecture, API gateway patterns, container orchestration (Kubernetes, Docker), autoscaling strategies, load balancing, resource scheduling algorithms, GPU computing, model serving optimizations, observability tools, API design, rate limiting, authentication/authorization
Preferred skills
Python, Rust, C++, Kafka, RabbitMQ, PostgreSQL, Redis, event-driven architectures, HPC cluster management, data pipelines, low-level systems programming, ML platform engineering, enterprise deployment
Technologies
vLLM, TGI, TensorRT-LLM, Kubernetes, Docker, Prometheus, Grafana, OpenTelemetry, Kafka, RabbitMQ, PostgreSQL, Redis
Responsibilities
Design and build core platform components for inference infrastructure; Optimize inference pipelines for latency, throughput, batching efficiency, and resource utilization; Build and debug production-grade code for real-time AI workloads; Implement intelligent routing systems for multi-model serving and canary deployments; Build high-performance telemetry and observability stacks for inference metrics; Conduct code reviews to maintain architectural consistency and production readiness
Seniority
Senior, hands-on IC