Application Software Engineer, Inference
Core
Design and optimize large-scale AI inference platforms to serve mission-critical models for SpaceX's launch vehicles and Starlink systems.
Role type
Senior IC application software engineer (LLM inference systems)
Builds
High-throughput, low-latency distributed inference infrastructure for internal SpaceX AI applications
Domain
Aerospace / Large Language Model Inference
Deliverable
production ML models
Required skills
Rust, C++, distributed systems design, low-level GPU optimization, model serving frameworks, system observability, CI/CD
Preferred skills
LLM inference engines (SGLang, vLLM, TensorRT-LLM), speculative decoding, agent SDKs, Kubernetes, gRPC, Python/Go
Technologies
SGLang, vLLM, TensorRT-LLM, Triton, PostgreSQL, ClickHouse, MongoDB, Docker, Kubernetes, gRPC
Responsibilities
Architect scalable distributed infrastructure for model serving; Optimize latency and throughput via GPU kernels and quantization; Build high-concurrency serving systems with 100% uptime; Own end-to-end components like request routing and SDK development; Benchmark and accelerate inference engines; Develop custom tracing and debugging tools; Create CI/CD infrastructure for deployment
Seniority
Senior, hands-on IC